The Generative AI Learning Penalty: Evidence from Chinese Secondary Education
Using 30 months of panel data on 26,811 Chinese students in grades 7-12, we study how generative AI affects homework productivity and learning. The data combine monthly closed-book exams, high-school and college entrance exams, and homework scores and completion time across nine subjects. We exploit staggered AI adoption in a difference-in-differences design. AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months. High-stakes entrance-exam scores fall by 18 and 24%, with the full penalty emerging only after about two years. The losses are largest in social science subjects, followed by STEM and languages, and are especially large for junior students, high-achieving students, and boys. The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores. AI users who maintain similar homework completion time as non-AI users experience small learning losses.
Given OPs axis definitions.
The max homework score for nonai is 115. But there are AI users scoring 115+? I’m confused why there are no better scores above 115 for nonai. Is that data insignificant? Are students with scores higher than 115 just being flagged as AI? It makes the data look unreliable.
I’d also really like to know if the data range that isn’t comparable 115+ is at all meaningful data. Because that’s where there is a major dropoff. There is obviously a dropoff of significant before 115. But, just from looking at how the data is being presented with this massive dropoff but with zero data for nonai.
I’d be interested in knowing what fraction of the total population is actually being represented in the 115+ dropoff. The way it’s presented it could literally be like 10 students.
This is not a defense of AI use. But more a criticism of how the data is being presented.
Also, Any student could be using AI as a resource similar to how I would use solutions manuals or previous tests to study back in my education. The good students aren’t gonna copy it verbatim and get flagged. Their going to get the right answer, and use it to learn the process so they can present their own work and actually learn the material.
AI is trash. But this is nothing different than before AI when students would copy the solutions manuals blindly or their peers work. Those students have always existed. AI just makes it slightly more accessible. But, really only slightly. And, honestly, Chegg was just as easy to copy back in the day.
Edit: This graph isn’t in the paper that OP linked. So, probably why it’s presented so badly. I’m assuming it’s from some click bait article then. There is a similar graph in the paper that is much more clear.
But the paper clarifies that this is an actual survey of students. So it’s not just from flagged AI homework. Which is good and bad in its own way.
The paper itself in section 5 supports what I hypothesized above. It’s not really about AI making students dumb. It’s about making it easier for students that already don’t want to learn to finish the homework. Same as the people that would copy solutions manuals.
Quote from section 5.
In this section, we show that the majority of AI students exhibit behavior consistent with homework outsourcing, spending little time completing homework assignments and experiencing learning losses. The remaining AI students spend as much time completing homework as non-AI students, receive higher homework scores, and have similar exam scores. In other words, they are learning as efficiently as the non-AI students.
The paper doesn’t argue it from what I read. But I’d also argue there is some bias in the survey from the questions themselves. Or at least how the paper lumps categories. The students scoring well in both exams and homework are less likely to consider their use of AI as meaningful enough to answer “used AI for homework”. It’s a problem with the survey. The survey is asking the students that actually learned the material to attribute their homework to AI alone in the same way the “copy and paste” AI users would.
The survey would probably benefit from having No AI, Some AI, All AI as it’s responses. Even in anonymous surveys people make these own interpretations of the questions in their head. And a student that used AI as a resource to learn is very likely to just choose “No AI” when presented with questions of how they got their final answers to homework. Because in their head they are thinking of the students that copy and pasted AI responses and think “I’m not like that”.
It’s likely why the double high scoring AI user sample is so low (paper says this). The survey is not allowing the response for this set of users to categorize themselves as using AI without feeling like they are like the copy and paste students. So those people are likely just self categorizing as “no ai” because the survey doesn’t allow them to distance themselves from the other population. Surveys are hard to write. Even ones where the users know it’s anonymous. Good students that use AI will see themselves as good students and the copy pasters as bad students. Human emotion plays a role and most people will not self categorize themselves further negatively than they feel they should be. They are more likely to select the imperfect category that is more positive.
Bad students don’t care though. They’ll admit to AI use. They already don’t care enough to learn the material. They have no reason to lie. So they’ll select “used AI”.
If you want a good survey you need very simple and minimal answer categories that the majority of the population will be able to self categorize themselves well. But, you don’t want to minimize to such a degree that your survey introduces self categorizing bias. I would argue that’s what happened here. The most important category is an extremely low sample size compared to the rest of the data.
And I’d also say it’s likely the majority of people in reality. They just self categorized as “no ai” because the survey didn’t allow them to categorize themselves more accurately without being associated with what they see as negative.
The entire paper seems to not address this strict categorization in its data that is likely introducing bad response data. But I’d have to look at the actual survey questions to know.
I think it’s a good paper that suffers from an overly strict self categorizing bias in its survey data.
Homeworks are useless, school bad, Einstein had bad grades (no he didn’t) crowd is very silent right now.
He had a bad grade in French language and literature

6 is the best grade in SwitzerlandLemmy is pretty unfriendly to anti-intellectualism, thankfully.
So… Below average kids should use AI? Is that what this says?
Looks like millennials / elder Gen z are going to be the most educated generation in history. Older generations were taught when they knew less about the world and education, younger generations are screwed both by the pandemic and now AI.
So Idiocracy was sort of right, only it’s not the eugenics “dumb people have more kids” reason they gave, it’s just us offloading our thinking to machines.
I’m sorry, but WTF is this chart without clear axes?

This is from the paper, and clearer to me.
Thank you. Spent entirely way too long figuring out WTF those axis were supposed to be.
EDIT: nevermind, I misread. I thought you were complaining about the one you posted… However, the image used by op is still a fairly obvious graph.
Edit 2: the axis of the graph posted by op are also clearly defined, vertical is exam scores, horizontal is homework. The definition shows above and below the graph itself.
It seems fairly obvious to me. It shows the clear mismatch between homework scores and test scores of AI users, where higher homework scores using AI correlate to lower test scores, clearly indicating that there is no learning, when compared to non-ai users, where higher homework scores directly translate to higher test scores as well.
Yeah I get it.
Looking at the OP graph again, I get it now, but initially I found the top label to be confusing. I thought it was part of the title, and that the Y axis was unlabeled.
Thank you.
Meanwhile on Lemmy: 3 downvotes.
I’m getting too old for this
Side tangent, but Lemmy is the worst with sources.
If I were admin, I’d make all users, new and extant, take an 8-minute course on clicking through reposts to finding original sources before they are allowed to submit posts.
That’s my rant.
I don’t think you can prevent this from being posted, though I agree that resource should be available. It’s just that our attention spans have evaporated, everything must be instant and with LLMs it’s only accelerating.
This is about having a brief second of critical thinking and then not upvote or downvote the low effort post and upvote the first critical thinking comment.
Another thing on Lemmy is that we have half a dozen “users” who don’t care about the content, they just automatically post high engagement content from other platforms and post it here (without declaring themselves as a bot account too). And the mindless scrollers (at least 60-70%) going by this post, just take the bait.
Not having karma didn’t help much in hindsight.
Just in case as it’s formatted a bit weirdly
X-Axis: Homework scores
Y-Axis: Exam scores
The implication being that if you use AI to do your homework and score highly on it, when it comes time to take the exam you’ll do very badly. If you don’t use AI to do your homework, your exam scores will more or less match your homework scores.
Still, the shapes of the curves are weird. I guess it suggests that students scoring around 100 on their homework are probably not relying on AI overly much, so their exam scores are still decently high. But, every student who has a higher homework score is relying on AI and, as a result, their exams suffer?
What really sucks about this is that I don’t think students have the self-discipline to not use AI. I’m old enough now to know that cutting corners to get good grades is dumb if it means I’m not learning. But, I wasn’t smart enough to know that back then. I absolutely would have cut corners to get those good grades and get through the homework more quickly. My parents wouldn’t have approved, but they weren’t technologically sophisticated enough to stop me.
That explain SO MUCH (meaning, the entire thing).
Damn, Chinese kids who don’t use AI are already putting in more than 110%.
My takeaway: letting AI do your homework for you is essentially the same as letting AI study for you.
Kind of obvious, but the data is still good to have.
Cognitive attrition.
Turns out you need to practice doing things. Who knew.
Hm, this is something that feels quite obvious but its very nice to have some hard data confirming it
Did AI come up with these axes?
This is actually a very good chart to explain outcomes. It shows that there’s a positive linear relationship between homework and exam scores for students that don’t use ai, and a terrible relationship for students that do use ai
I was taught to always label your axes.
They are labelled. It’s just the Y label is next to the title.
Which is the Y label?
“China* average exam score”
Wouldn’t a bar graph be better for this sort of data? The line keeps suggesting change over time to me and it’s making the whole thing hard to read, even with OP’s note.
Disagree-- line graphs are good for most continuous x axes, not just time. There’s no reason to group the homework scores but not the exam scores, as most usages of a bar graph here would, or to use a bar graoh with 1 bar per possible hw score (which is basically just a line chart anyways).
The issue with the chart is just that the axes aren’t labelled well at all. X is hw, Y is exam, and it looks like rather than raw score it’s using… score divided by the average score, as a %? So how much better or worse the score is than the average, as a percentage?
It got me confused too at first but once you think about it it makes a lot of sense : the gray line is mostly straight, efforts get rewarded, but the more they use AI in homework, the worst they do on exams, even worst than the worst of the no AI crowd
Funny is that when I saw the graph I immediately thought that most people would be confused.
Home work and learning have never been about getting the right answers
You’re right, homework was invented as punishment
Promotion at work also isn’t about being the best at your job.
Between social requirements and gaming metrics, it’s a wonder we haven’t collapsed yet.
Why are boys more susceptible?
Given the overlap with “high achievers” and China’s cultural focus on men/male children, I’d imagine it’s an emergent property of widespread gender bias rather than intrinsic to AMAB academic capabilities.
I guess I’m not clear on if the study is saying
boys use AI more
or that
their scores suffer more when they use it.
This is conjecture because I can’t access the full paper at this time, but based on:
The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores.
I’m guessing this is an artifact of “students under higher pressure to perform are more likely to use AI more thereby lowering their scores”. Given Chinese patriarchical cultural biases, I’d imagine boys are under higher pressure on average, thereby leading to greater usage.
So perhaps it’s actually the second one, as the abatract is unclear if they’re studying performance relating to usage or not. Without checking methodology, it’s unclear if they’re controlling for usage.
From other things I’ve read, men in general are more likely to use AI, not just in China. So I would assume the first.
I suspect it’s the second but now that you mention it, it is a bit ambiguous. I’m assuming the study would control for usage to compare outcomes across similar usage metric cohorts.












