Key takeaways
Hand-scoring 4,000 essays costs roughly 530 hours of teacher time — about $22,000, scored just once.
Rater drift is real: scorers get more generous and give less feedback by the end of a long stack.
One centrally calibrated AI assistant keeps every classroom scoring against the same standard.
Locking parameters before rollout made adoption easy — zero office-hours attendance needed.
Best for: district leaders running a writing benchmark who need consistent, fast scoring district-wide.
Owensboro Public Schools has been running district-wide writing benchmarks since 2019. Three times a year, every student in grades 3 through 11 sits down, reads a set of sources, and writes an argumentative response on demand. Amy Bellamy, the district's Instructional Coordinator for ELA and Social Studies, describes it simply: after Kentucky pulled on-demand writing out of the state test, "you have to have some sort of system for writing. So this is our system."
That decision is about to matter more, not less. Due to legislative changes through HB 257, the 2026-2027 KSA will no longer test writing and editing and mechanics. Districts that want to know whether their students can write will have to find out for themselves. Owensboro figured that out seven years ago.
This fall they ran their first benchmark on EnlightenAI. Close to 4,000 pieces of student writing came through in a single two-and-a-half-week window. Here is what that actually involves, what it used to cost, and what changed.
People hear "benchmark" and typically picture a multiple-choice exam. A writing benchmark is a different animal.
Every student gets the same prompt and the same sources for their grade level. They write a full response, typically several paragraphs, arguing a position and citing evidence. Then every one of those responses has to be read by a human, scored against a multi-trait rubric (claim, evidence and sourcing, counterclaim, organization, conventions), and turned into feedback the student can use and data the district can act on.
Owensboro's first window ran August 31 through September 17. Third graders wrote on iPads. Sixth graders and up wrote on Chromebooks. Everyone got one submission, no second attempts, because the point of the first benchmark of the year is diagnostic: "it's just to kind of see where we are," as Amy put it. Two more windows follow in November and February so the district can measure growth against that baseline.
The math on 4,000 essays is daunting.
A teacher scoring a multi-paragraph argumentative essay against a five-trait rubric, and writing even brief feedback, needs somewhere around eight minutes per response (some will take longer, others shorter, but for illustrative purposes let’s settle on eight).
4,000 essays × 8 minutes = 32,000 minutes, or about 530 hours of scoring.
That’s more than thirteen 40-hour workweeks, or roughly $22,000 worth of teacher time at an estimated hourly rate of $41, based on Kentucky’s average teacher salary. Spread across a district's ELA teachers, it lands as several evenings and a weekend for each of them, three times a year, on top of the writing they already assign and grade in their own classrooms. And that number assumes single scoring. Any district that wants its benchmark data to be trustworthy across schools would ideally double-score at least a sample, which pushes the total higher.
So most districts make one of two compromises. They score slowly, and results arrive weeks after students wrote, or they score fast, and consistency suffers.
EnlightenAI helps teachers deliver instant, rubric-aligned AI writing feedback so students can practice, revise, and improve faster. It's a simple way to start grading essays more efficiently.
Hand-scoring 4,000 essays is not just time-intensive. It is hard to execute reliably in ways that are well documented.
Researchers have long studied "rater drift," the tendency for a scorer's standards to shift over the course of a session. In one study of raters each scoring 28 composition tasks, the scores assigned to the first 16 tasks were significantly lower than those assigned to the last 12, and the last four tasks were scored highest. The same study found the frequency of comments given by raters on the first 12 essays was significantly higher than on the last 16. Tired scorers get more generous and say less. The student whose essay lands at the bottom of the pile gets a different score, and less feedback, than the student at the top.
Then there is the gap between scorers. The standard the testing industry uses to judge whether two raters agree is a statistic called quadratic weighted kappa (QWK). Under the widely cited Williamson, Xi, and Breyer framework, agreement between automated and human scoring must reach at least .70 to be considered acceptable. Two teachers scoring the same stack often don't clear that bar. In a study with DREAM Charter Schools, two teachers scoring the same essays on the same scoring guide agreed at a QWK of 0.52. EnlightenAI's TA, trained by one of those teachers, agreed with that teacher at 0.77. The TA was more consistent with the teacher who trained it than a second expert teacher was.
That’s what makes Owensboro’s commitment to writing benchmarks remarkable. Running a district-wide writing benchmark three times a year is a herculean undertaking: coordinating across schools, aligning teachers on scoring, and making time to read thousands of student responses on a deadline. Sustaining that work since 2019 takes tremendous effort and a shared conviction that understanding how students write is worth it.
When Owensboro evaluated EnlightenAI, this was the thing Amy named first: "The main things we were excited about with Enlighten is the ability to have that admin dashboard, be able to see school-level data and district-wide data."
Owensboro's rollout of EnlightenAI worked because the district team treated the benchmark as a design problem and made three decisions before a single student logged in.
They calibrated the TA centrally, once
Rather than asking 100-plus teachers to each train their own scoring assistant, Amy's team built one assignment per grade level: the district prompt, their district rubric, and high, medium, and low anchor papers that a committee of teachers reviewed. The instruction for a top score on sourcing, for instance, was explicit: cite both sources. Every teacher in the district received the same assignment, complete with a calibrated TA. As Amy said, "it needs to be trained the same way for the benchmark." One standard, every classroom.
They locked the parameters
One submission per student. Feedback format set by grade band: for elementary, two glows and two grows in under 50 words; for secondary, feedback by rubric trait. No teacher-level customization for the benchmark. Teachers still review and approve every score before students see it, but the standard doesn't drift from room to room.
They made the process dead-simple for staff
A student-facing walkthrough. Step-by-step directions for teachers (select the district benchmark, assign it to your Clever classes, share the link). A short tutorial for principals and instructional coaches on pulling reports. Amy even asked for the teacher guide in editable form so she could update the assignment names herself each cycle. Different people needed different things, and the district said so up front.
The EnlightenAI team had scheduled three office-hours sessions during the window to help teachers through it. Nobody showed up. Teachers were already running the benchmark on their own, and the few questions that came up were easily handled by the district team.
Under hand-scoring, feedback reaches students and results reach leaders weeks after the writing happens, which can be a month or more for a full district. For Owensboro on this benchmark, as submissions came in, the TA scored them against the calibrated rubric in seconds, teachers reviewed and released them, and students had rubric-level feedback while the prompt was still fresh.
At the district level, Amy was looking at live results before the window closed. On August 31, the first day of the window, the dashboard already showed 398 submissions across grades 6 through 8, and an early trend was already visible in seventh grade: students were strong on sourcing and struggling on counterclaims. That is a reteaching objective, identified on day one, with two and a half weeks of window still to go.
The research on why this matters is not ambiguous. A meta-analysis of formative writing assessment in grades 1 through 8 found that feedback to students about their writing from adults enhanced writing quality with an average weighted effect size of 0.87, and formative assessment overall carried a weighted mean effect size of .61. Effects that size are rare in education research. But they depend on feedback actually reaching the student in time to act on it. A score that arrives in November for an essay written in September dilutes the impact.
Owensboro can now see how student writing is developing across the district, at each campus, and in each classroom. They can see which rubric traits students are strongest on and which are the highest-leverage growth areas, and put their coaching and PD where the writing says it's needed. And they'll be able to measure whether it worked when the November window opens.
"The feedback on the benchmark writing assignments was actionable and specific. The time saved by using this system allowed our teachers more time to analyze the data and plan targeted instructional next steps. We look forward to continuing our partnership with EnlightenAI and seeing our students' writing skills grow."
- Amy Bellamy, Instructional Coordinator for ELA and Social Studies, Owensboro Public Schools
Five hundred hours of grading time returned to teachers as time to plan, all while speeding up turnaround times for scoring and feedback while boosting consistency. That is the story of Owensboro’s first benchmark with EnlightenAI.


