GPT-BERT Wins BabyLM 2024, Beating Llama 2 70B With 100 Million Words
Updated
Updated · MIT Technology Review · Aug 24
GPT-BERT Wins BabyLM 2024, Beating Llama 2 70B With 100 Million Words
1 articles · Updated · MIT Technology Review · Aug 24
Summary
GPT-BERT won BabyLM 2024 after pretraining on about 100 million words and still beating Meta’s Llama 2 70B on one grammar benchmark.
That result targets AI’s data-efficiency gap: toddlers start producing grammatical sentences after roughly 10 million to 30 million words, while frontier models train on vastly larger corpora.
BabyLM was built to test whether child-scale training can produce strong language models, but organizers say popular baby-like ideas such as curriculum learning have not worked as well as expected.
Researchers increasingly think closing the gap may require more than text—adding vision, interaction and social learning—though multimodal and interactive baby-inspired models still lag standard approaches.
The broader payoff could extend beyond English chatbots, helping universities and smaller-language communities build capable models with limited data while giving cognitive scientists a new tool to study human language learning.