The "distillation" argument always seemed pretty weak to me. If one argues that training an AI on copyrighted content is merely analogous to reading a book (and therefore is fully transformative), then it stands to reason that training an AI on other AI outputs would also be fully transformative.
It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.
I don't think they're accusing Chinese organizations of violating copyright or patent treaties, they're more so observing undesired behavior and pushing for domestic strategies to prevent it from happening.
There was a fairly convincing discussion on HN about a month or so ago that argued that this is enabled by the large difference between the monthly subscription costs and the API costs to access the very same models. That the primary purpose is not distillation but pricing arbitrage. Data collection for distillation is just a bonus side effect.
Bringing the prices of monthly subscriptions much closer to API rates would kill the arbitrage business and make distillation a much more expensive (and harder to hide) activity. Raise one, lower the other - whatever.
It mostly involves sorting out Python dependency version conflicts.
Also, it takes a lot of compute. It still involves training a model, which for a very small model requires a GPU with several times more ram than the final model size, which will need to run for days to weeks. Usually you're also running the model that is being distilled from, but in this case it would require a huge database of chat logs, instead.
The "distillation" argument always seemed pretty weak to me. If one argues that training an AI on copyrighted content is merely analogous to reading a book (and therefore is fully transformative), then it stands to reason that training an AI on other AI outputs would also be fully transformative.
It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.
I don't think they're accusing Chinese organizations of violating copyright or patent treaties, they're more so observing undesired behavior and pushing for domestic strategies to prevent it from happening.
This argument only makes sense if you equate a printed book with a chatbot. They aren't the same thing, and the chatbot's outputs aren't copyrighted.
There was a fairly convincing discussion on HN about a month or so ago that argued that this is enabled by the large difference between the monthly subscription costs and the API costs to access the very same models. That the primary purpose is not distillation but pricing arbitrage. Data collection for distillation is just a bonus side effect.
Bringing the prices of monthly subscriptions much closer to API rates would kill the arbitrage business and make distillation a much more expensive (and harder to hide) activity. Raise one, lower the other - whatever.
How does distillation actually work? If I wanted to distill a model and I have enough money and compute, what’s the process?
It mostly involves sorting out Python dependency version conflicts.
Also, it takes a lot of compute. It still involves training a model, which for a very small model requires a GPU with several times more ram than the final model size, which will need to run for days to weeks. Usually you're also running the model that is being distilled from, but in this case it would require a huge database of chat logs, instead.
So China can't steal the data that US companies stole fair and square?
https://www.techopedia.com/trump-administration-openai-train...