How much of IELTS examiner judgment do you think is actually reproducible by a model
I’ve spent two years building AI scoring for IELTS, and I still don’t have a clean answer to this.
Here’s where I’ve landed so far. Some parts of the rubric are mechanical and a model handles them well: word count, task coverage, paragraph structure, grammatical range, whether the response actually answers the prompt. Reading and Listening are trivially checkable — there’s a key.
Then there’s the part that isn’t. IELTS Writing has a criterion around lexical sophistication, and Speaking has one around natural fluency. Human examiners disagree with each other on these — published inter-rater agreement isn’t as tight as people assume. So when my model outputs 7.0 on “Lexical Resource,” I genuinely don’t know if it’s right or just confidently averaging.
What I’ve observed from ~1,200 essays graded through an earlier version: the model skews generous on Coherence and harsh on Grammar. It rewards connective words as if they were coherence, when a real examiner reads them as mechanical when overused.
Two things I’d like to hear:
— If you teach IELTS or have been an examiner: which criterion do you think is least automatable, and why?
— If you’ve used any AI scoring tool as a test taker: where did the score feel wrong, and did you find out later what your real band was?
I’m launching Tuesday, so partly I’m asking because I want to know what to caveat honestly rather than overclaim.

Replies