GPT-4o boosts grades without building skills, Bocconi study finds

A Bocconi University study of over 1,000 students found GPT-4o improved marketing assignment grades by nearly a full letter grade, yet revealed a critical gap: performance gains did not correlate with actual learning. The finding underscores an emerging tension in AI-assisted education where models excel at mimicking surface-level competencies that grading rubrics reward, while independent reasoning atrophies. This pattern has broader implications for workforce readiness and institutional credibility as employers increasingly question whether credentials earned with AI assistance signal genuine capability or merely credential inflation.
Modelwire context
Analyst takeThe Bocconi finding isolates a specific failure mode: AI doesn't just help students produce better work, it rewards them for skipping the cognitive steps that rubrics don't explicitly grade. The gap between output quality and learning retention suggests institutions are optimizing for the wrong signal.
This connects directly to the August study on AI agents' miscalibration of their own performance. Just as coding assistants inflate confidence by 20 percentage points while misjudging constraints, students using GPT-4o are receiving inflated grades that don't reflect actual capability. Both point to the same underlying problem: systems that game surface metrics while internal competence erodes. The difference is scope. Miscalibrated code agents create production failures in isolated systems; miscalibrated credentials create systemic hiring and promotion failures across entire cohorts entering the workforce.
If employers begin screening for 'AI-era credentials' with separate assessment protocols within 18 months, or if any major university implements dual grading (one for AI-assisted, one for independent work), that signals institutions are treating this as a market failure requiring structural response rather than a temporary adjustment problem.
Coverage we drew on
- AI agents have no sense of time and are not aware of it · The Decoder
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGPT-4o · Bocconi University · OpenAI
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “The skills that earn top grades are the ones AI can fake best”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.