Fine tuning is not necessarily oriented at making a model better at generic "intelligence" tasks. In fact, the goal is often specifially to improve model performance at narrow tasks while recognizing it will get worse at other tasks.
I clicked through the leaderboard. It seemingly isn't sorted by best or worst performing, but rather by absolute change in performance - improved and worsened models are interleaved. Regardless, I selected the model that worsened the most, and investigated what they claimed their fine tuning accomplished; in fact they claimed nothing, their model sheet was the default trl model sheet without even the training set listed. The second worst performing was the same, but they at least listed their training set. It was a set of millions of word math problems. So we would expect the model to get better at those problems and worse at everything else.
On the other hand, the most improved model on the table had an actual model sheet, where they claim the model has been trained to be "more helpful". They claim that by ablating model refusal the model ends up doing better on benchmarks as well. That claim seems to have been borne out.
Obviously if you train a model for some purpose it will not necessarily do well at that purpose. But I would expect popularily used models to be ones that actually do what they say they do. I don't think people need to be told that sometimes models are badly trained. And the claim that this project actually has data to support, that finetuning on one task does not necessarily improve performance on another task, is so obvious that it doesn't need to be stated.
Fine tuning is not necessarily oriented at making a model better at generic "intelligence" tasks. In fact, the goal is often specifially to improve model performance at narrow tasks while recognizing it will get worse at other tasks.
I clicked through the leaderboard. It seemingly isn't sorted by best or worst performing, but rather by absolute change in performance - improved and worsened models are interleaved. Regardless, I selected the model that worsened the most, and investigated what they claimed their fine tuning accomplished; in fact they claimed nothing, their model sheet was the default trl model sheet without even the training set listed. The second worst performing was the same, but they at least listed their training set. It was a set of millions of word math problems. So we would expect the model to get better at those problems and worse at everything else.
On the other hand, the most improved model on the table had an actual model sheet, where they claim the model has been trained to be "more helpful". They claim that by ablating model refusal the model ends up doing better on benchmarks as well. That claim seems to have been borne out.
Obviously if you train a model for some purpose it will not necessarily do well at that purpose. But I would expect popularily used models to be ones that actually do what they say they do. I don't think people need to be told that sometimes models are badly trained. And the claim that this project actually has data to support, that finetuning on one task does not necessarily improve performance on another task, is so obvious that it doesn't need to be stated.
[flagged]
[flagged]