What Drives Assessment Reliability? Findings from a Large-Scale Assessment Analysis
Assessment teams often ask a familiar question: What can we do to improve reliability?
The traditional answer is straightforward—add more items. While test length is known to influence reliability, assessment programs frequently face practical constraints such as limited seat time, development costs, and learner fatigue.
To better understand what drives reliability in operational assessments, our team conducted a research study examining more than 80 assessments administered within a professional learning environment.
Our goal was to identify which assessment characteristics were most strongly associated with reliability and to explore how quantitative and qualitative indicators of assessment quality work together.
The Study
For each assessment, we collected a variety of psychometric and content-related indicators, including:
- Reliability (Cronbach's alpha)
- Average item discrimination
- Percentage of items with low discrimination
- Average item difficulty
- Average guessing parameter
- Number of items on the assessment
- Score variability
- Qualitative item review scores
- Differential item functioning (DIF) indicators
The qualitative review process evaluated items against established item-writing principles, including clarity of the stem, response option quality, ambiguity, parallelism, clueing, and other characteristics that can affect the testing experience.
We then conducted correlation and regression analyses to identify which variables were most strongly related to reliability.
What We Found
Several expected patterns emerged.
Reliability increased as assessments contained more items and as score variability increased. These findings are consistent with well-established psychometric theory.
However, one result stood out above all others.
Item Discrimination Was the Strongest Predictor of Reliability
Average item discrimination demonstrated an exceptionally strong relationship with reliability.
Across the assessments included in the study, the correlation between average item discrimination and reliability was approximately 0.93.
To further investigate this relationship, we fit a regression model using average item discrimination as the sole predictor of reliability.
The results showed that average item discrimination alone explained approximately 86% of the variability in reliability across assessments.
In practical terms, assessments containing items that effectively differentiate between more knowledgeable and less knowledgeable learners were substantially more likely to produce reliable scores.
Test Length Remains Important
We then expanded the model to include both average item discrimination and the total number of items on the assessment.
Together, these two variables explained approximately 97% of the variability in reliability.
This finding reinforces two important principles:
- Longer assessments generally produce higher reliability.
- Item quality, as reflected by discrimination statistics, remains critically important even after accounting for assessment length.
What About Qualitative Item Reviews?
While statistical indicators emerged as the strongest predictors of reliability in our analysis, the findings also reinforced the important role of qualitative item reviews in the assessment development process.
Our review framework evaluated items against established item-writing principles, including clarity of the stem, response option quality, ambiguity, parallelism, unnecessary complexity, and potential clues to the correct answer. These reviews provide valuable evidence that cannot be obtained from psychometric statistics alone.
Reliability tells us how consistently an assessment measures performance. Qualitative reviews help ensure that items are clear, fair, and aligned with the intended learning objectives. In other words, statistical analyses help us understand how well items perform, while qualitative reviews help us understand why they may perform the way they do.
Although qualitative review scores were not among the strongest predictors of reliability in this study, they remain a critical component of a comprehensive quality assurance process. By identifying potential item-writing issues before administration, qualitative reviews help support content validity, improve the learner experience, and strengthen the overall defensibility of assessment results.
The findings suggest that the most effective assessment programs leverage both approaches. Psychometric analyses provide evidence of item and test performance, while qualitative reviews provide expert insight into content quality and opportunities for improvement. Together, these complementary sources of evidence support the development of assessments that are both technically sound and instructionally meaningful.
Implications for Assessment Programs
The findings suggest several practical recommendations for assessment programs:
- Monitor item discrimination statistics routinely.
- Revise or replace items that consistently show weak discrimination.
- Continue conducting structured qualitative item reviews.
- Use statistical evidence and content review findings together when making assessment decisions.
- Consider item quality alongside test length when evaluating reliability.
Most importantly, organizations should recognize that reliability is only one dimension of assessment quality. High-quality assessments require both strong psychometric performance and sound content design.
What These Findings Mean for Assessment Programs
Our analysis demonstrated that item discrimination is one of the strongest drivers of assessment reliability, while assessment length continues to play an important supporting role.
At the same time, the study reinforced the value of qualitative item reviews. Although qualitative review results were not the strongest predictors of reliability, they provide essential evidence about clarity, fairness, and content quality that cannot be captured through statistics alone.
The most effective assessment programs do not rely on a single source of evidence. They combine expert review, psychometric analysis, and continuous improvement to build assessments that are both reliable and defensible.










