Stanford Professor Has Students Build 24 AI Evals in 1 Class to Test Model Power
Updated
Updated · Financial Times · Jul 24
Stanford Professor Has Students Build 24 AI Evals in 1 Class to Test Model Power
3 articles · Updated · Financial Times · Jul 24
Summary
Three hours of class time were enough for Stanford undergraduates with no coding background to build working AI evals and webpages comparing models against their own criteria.
The exercise taught students to test model output systematically rather than use AI passively, giving them a way to hold systems from companies such as OpenAI and Anthropic accountable.
Projects ranked models across 24 criteria, including Brazilian election accuracy, resistance to user pressure, and whether answers to culturally specific questions were stronger in English than Burmese.
The professor argues universities should pair AI-free courses with classes that teach evaluation skills, because students need subject expertise to design useful tests and many existing benchmarks are flawed.
Next year he plans to extend the approach to executives and MBAs, framing AI evals as a way for users to keep control as power concentrates in a handful of AI companies.
With AI evaluation costs soaring, can critical AI literacy ever be truly democratized beyond elite institutions?
Does teaching students to command AI build true critical thinking, or just a more sophisticated form of technological dependency?
When private AI 'evals' become valuable IP, are we training citizens or just highly skilled labor for Big Tech?
Critical AI Literacy and Governance: Stanford’s Free Systems Lab and the Urgent Need for Evidence-Based AI Education
Overview
Stanford’s educational philosophy centers on empowering individuals to engage critically with artificial intelligence, ensuring that AI enhances human agency and supports democratic values. The Free Systems lab, established in 2023 and led by Andrew Hall, puts this vision into practice by exploring how AI can positively transform politics and be overseen by people, not just adopted as tools. This approach highlights the importance of moving beyond using AI to understanding its societal impacts, including the need for strong governance and awareness of AI’s human-like behaviors. Cultivating critical AI literacy is a key principle, preparing students to thoughtfully navigate and shape the future of technology.