When Reading Benchmarks Hurt Multilingual Learners
A benchmark is only useful when we understand what it tells us—and what it doesn't.
Good leaders hold every student to high expectations.
Great leaders know that high expectations and identical benchmarks are not the same thing.

What If the Student Isn't the Problem?
A multilingual learner takes the fall reading screener.
The score comes back below benchmark.
Red.
Almost immediately, a story begins to form.
The student is struggling.
Maybe they're missing foundational skills. Maybe they need fluency or comprehension work. Most likely they need a reading intervention.
Sometimes those conclusions are right.
But systems thinkers have learned to ask a different question:
What if the student isn't the problem?
What if the problem is the benchmark?
Over the past several weeks, I've written about designing equitable MTSS systems for multilingual learners. We've talked about doing no harm, interpreting assessments appropriately, examining growth over time, strengthening Tier I, and using tools like the Reading/Language Matrix to put academic performance into the context of language development.
All of those ideas lead to a potentially uncomfortable conclusion:
Sometimes our benchmarks cause us to see a problem that isn't actually there.
And what happens next can have real consequences for students.
Below Benchmark Is a Description, Not a Diagnosis
Let's start with something simple.
Below benchmark describes performance. It doesn't diagnose the cause.
A reading screener tells us how a student performed on specific reading tasks at a specific point in time. For multilingual learners, that performance can be influenced by many factors: English proficiency, vocabulary, background knowledge, previous educational experiences, literacy in another language, quality of Tier I instruction, and how long they've been learning English.
That's why the appropriate response to a low screening score should be:
"We need to look closer."
Not:
"This student needs intervention."
That distinction matters, because once we identify a student as struggling, we begin making decisions based on that assumption. We provide intervention. We change schedules. We may remove the student from other instruction.
And if the original assumption was wrong, the solution probably will be too.
But Aren't These Assessments Valid for Multilingual Learners?
This is where the conversation gets more complicated.
I often hear some version of this:
"But multilingual learners were included when these assessments were developed. The benchmark is still valid."
Let's use DIBELS as an example.
DIBELS is a valuable screening tool, and its benchmark goals aren't arbitrary numbers. DIBELS describes its benchmark cut scores as predictors of whether students are likely to meet later reading proficiency goals. Its technical documentation explains that the scores were developed by examining how well DIBELS performance predicted later reading outcomes.
That's useful information.
But here's the distinction I think we sometimes miss:
Predicting a future problem isn't the same as diagnosing its cause.
Imagine multilingual learners who haven't consistently received linguistically appropriate Tier I instruction. They may very well be at greater risk of later reading difficulties.
The screener may correctly identify that risk, but not what created it.
Is there an underlying reading difficulty?
Is the student still developing the English vocabulary and oral language needed to demonstrate what they know?
Have they had sufficient opportunities to connect new English words with concepts they already know?
Has Tier I instruction been designed to support second-language development?
The assessment can't answer those questions for us.
In fact, DIBELS itself distinguishes benchmark screening from diagnosis and describes its measures as tools for identifying students who may need additional instructional support and for monitoring response to instruction.
The data may accurately predict a future problem while we inaccurately diagnose its cause.
That's a systems-thinking problem.
High Expectations Don't Require Identical Benchmarks
At this point, someone usually raises an important concern.
"Aren't we lowering expectations for multilingual learners?"
No.
We should have incredibly high expectations for multilingual learners.
But high expectations and identical expectations at every point in a student's learning trajectory are not the same thing.
Research on English learners has long suggested that developing the social and academic English proficiency necessary for success in school commonly takes years; often an average trajectory of approximately five to seven years.
Think about what that means.
If we know students are progressing through a multiyear process of developing English proficiency, why would we expect their performance on standardized English reading measures to follow exactly the same trajectory as peers who already possess that English proficiency?
That's not rigor.
It's ignoring what we know about language development.
This doesn't mean changing the destination.
It means understanding the trajectory.
Classroom Expectations Are Different
This distinction becomes especially important when we compare standardized assessments with classroom learning.
I absolutely want multilingual learners engaged with grade-level content. In the classroom, a skilled teacher can scaffold access to that content. They can use visuals, modeling, strategic grouping, explicit vocabulary instruction, background knowledge, home-language resources, sentence supports, and other strategies that allow students to engage with rigorous learning while they continue developing English.
The learning target remains ambitious.
A standardized reading screener is different.
Many of those supports aren't present.
So when we interpret the resulting score, we need to remember that the student isn't only demonstrating reading performance.
They're demonstrating reading performance through their current level of English.
That context matters.
Maybe Growth Is the Better Question
This is why I've become increasingly interested in expected growth rather than simply static benchmarks.
Response to Intervention has a really important word in the middle:
Response.
We should be asking:
How is this student responding to instruction?
Suppose a multilingual learner begins the year significantly below a standardized benchmark but makes tremendous reading and language growth over the next several months.
The student is likely still below benchmark.
But are they struggling?
That's a very different question.
The Reading/Language Matrix I shared in my last post is one way we've tried to create better context for those conversations. It allows teams to consider reading performance alongside oral language proficiency rather than treating the reading score as an isolated data point.
Instead of only asking:
"Did the student meet the benchmark?"
we can ask:
"Given where this student started academically and linguistically, are they making the growth we should expect?"
That doesn't lower expectations.
I'd argue it creates much more meaningful ones.
When the Benchmark Starts a Harmful Cycle
This takes us back to the commitment I introduced two blogs ago:
Do no harm.
Remember the vacation activity?
I asked you to describe your favorite vacation without using a single word containing the letter N. You probably looked like a much weaker writer than you actually are. Now imagine I used that performance to identify you as struggling.
Then I pulled you from another important class to provide an intervention designed to improve a skill that wasn't actually your problem. You continued to struggle. What would I conclude?
Probably that you needed more intervention.
That's how a poorly interpreted benchmark can start a harmful cycle:
Below benchmark
↓
Student identified as struggling
↓
Intervention doesn't address the actual need
↓
Student loses time from Tier I or language development
↓
Performance remains low
↓
Original assumption appears confirmed
↓
More intensive intervention
At every step, the adults may be trying to help.
But good intentions don't protect students from poorly designed systems.
The Cost Goes Beyond Lost Instruction
There's another consequence that's harder to see. Think again about how you felt during that vacation activity.
Imagine experiencing that feeling every day—and repeatedly being shown evidence that you aren't meeting expectations.
How long before you start believing:
"I'm just not good at reading."
Student self-efficacy begins to erode.
Now imagine you're the teacher.
You repeatedly see multilingual learners failing to meet the same benchmarks as their English-proficient peers.
Eventually, your expectations can change too.
"These kids just aren't ready."
"They don't have the background knowledge."
"They always struggle with reading."
And when those beliefs spread across a faculty, something even more damaging can happen:
We stop believing that what we do can substantially change the outcome.
That's a collective efficacy problem.
This is why inappropriate expectations aren't harmless.
Benchmarks don't just influence how we see data.
They can influence how students see themselves—and how teachers see students.
And Sometimes the Cycle Goes Even Further
When intervention doesn't produce the expected result, schools naturally begin asking whether something else is happening.
That can eventually lead toward special education referral.
To be clear, multilingual learners can absolutely have learning disabilities, and students who need special education services should receive them.
The problem occurs when language development, inadequate opportunity to learn, or inappropriate instruction is repeatedly interpreted as evidence of disability.
Consider the system we've just described:
Weak linguistically responsive Tier I instruction.
↓
An English reading score below benchmark.
↓
An intervention based on an incomplete diagnosis.
↓
Limited response because the intervention doesn't address the actual need.
↓
Increasing concern.
↓
Possible disability referral.
No single person intended that outcome.
The interaction among the systems produced it.
That's exactly why systems thinking matters.
Before You Say "This Student Is Behind"
The next time a multilingual learner scores below a reading benchmark, don't ignore the score, but don't let the score finish the conversation either.
Ask five questions:
What does this benchmark actually tell us—and what doesn't it tell us?
What is this student's current English-language proficiency and language-development trajectory?
Is the student making appropriate academic and language growth from their starting point?
Has the student had sustained access to high-quality, linguistically appropriate Tier I instruction and English language development?
Does additional diagnostic evidence point toward a reading need, a language need, both—or a systems issue?
That's not lowering expectations.
That's improving our diagnosis.
Keep the Destination. Fix the Measurement.
I don't want lower expectations for multilingual learners.
I want better ones.
Expectations that recognize where students are starting, demand meaningful growth, maintain access to rigorous grade-level learning, and intentionally move students toward the same ambitious destination.
A static benchmark can be an important piece of information.
It just shouldn't become the entire story.
High expectations aren't defined by giving every student the same benchmark at every point in their journey.
They're defined by believing every student can grow—and designing a system capable of getting them there.
Reflection Question
Think about a multilingual learner your team identified as "below benchmark" last year.
What evidence demonstrated that the student had a reading problem rather than a language-development, opportunity-to-learn, or systems problem?




Comments