Plan well before introducing separate question sets for DU admission
Dhaka University is considering an important reform: introducing separate Science Unit (Ka unit) question sets for Bangla-medium and English-medium applicants so that each group is assessed in a way that better reflects its curriculum. Bangla-medium students, and of course madrasa students, study quite different books and syllabi from their English-medium peers. As a result, the latter often face problems studying for most public university admission tests. Under this situation, DU’s intention appears inclusive and deserves appreciation. Recognising differences in educational backgrounds is consistent with widening access to higher education.
However, inclusiveness alone is not sufficient. For a high-stakes admission examination, inclusiveness must be accompanied by scientific evidence that the assessment remains fair to everyone. As a Professor of Psychology and a researcher in psychometrics, I am not arguing either for or against separate question sets. My concern is much simpler: whatever system is adopted must be demonstrably valid, reliable, comparable, and fair.
The first scientific question is whether the two examinations would measure the same underlying ability. If different question sets assess different knowledge or cognitive skills, no statistical method can fully make their scores comparable afterwards. Establishing construct equivalence, meaning that the assessment would measure the same thing and have the same meaning for test takers of different groups, is therefore the essential first step.
The second question concerns merit ranking. If separate question sets are used, will applicants be ranked separately by curriculum, or will their scores be placed onto one common scale and combined into a single merit list? Either approach is possible, but whichever is chosen should be supported by a transparent scientific justification.
Modern psychometrics has addressed similar challenges for decades. International examinations such as the SAT, GRE, and TOEFL routinely administer multiple versions of the same examination while maintaining comparable scores. This is achieved through carefully designed psychometric procedures, including Item Response Theory (IRT), test equating, common anchor items, differential item functioning (DIF) analyses, and measurement invariance studies.
The important point is that separate question sets are not inherently fair or unfair. The scientific question is whether scores from those sets can measure the same underlying ability with comparable accuracy. If that evidence exists, separate question sets can improve inclusiveness without compromising fairness. Otherwise, the reform may unintentionally undermine public confidence.
For MCQ-based examinations, one practical approach is to include a carefully selected set of common anchor items—same questions—in both versions of the examination. These common questions provide the statistical link needed to compare the relative difficulty of the two forms. Depending on the nature of the data, appropriate IRT models—such as the Rasch, two-parameter logistic (2PL) or three-parameter logistic (3PL) model—to determine how difficult each question is and how well it measures ability.
The next step is test equating, through which raw scores from different question sets are converted onto a common reporting scale. The statistical calculations themselves are relatively straightforward once the test results are available. The real challenge is ensuring that the different question sets were designed in advance to be comparable by following the same test blueprint and including common anchor items.
Equally important is examining each item for Differential Item Functioning (DIF) to ensure that no question unintentionally favours one curriculum over another. Measurement invariance analyses can further determine whether identical scores have the same meaning across different educational backgrounds. If written components are included, additional safeguards become necessary, including detailed scoring rubrics, assessor training and evidence of inter-rater reliability.
If Dhaka University decides to proceed with this proposal, several practical steps would strengthen confidence in the reform. These include conducting a pilot study before full implementation; establishing a calibrated item bank, in which the statistical properties of each question have already been measured and documented; designing examinations with common anchor items. Also, preparing a psychometric team to conduct equating analyses; publishing a technical report summarising reliability, validity and equating results; and seeking advice from an independent panel of psychometricians, educational measurement specialists, statisticians and subject experts should be carried out.
Admission examinations shape educational opportunity for thousands of talented young people every year. Decisions of this importance should not rest solely on good intentions or public opinion. They should be supported by scientific evidence that the assessment measures what it intends to measure and does so fairly for every applicant. Dhaka University has an opportunity to demonstrate that inclusiveness and fairness are not competing values. With appropriate psychometric planning and transparent reporting, it can achieve both.
Dr Md. Kamal Uddin is professor in the Department of Psychology at Dhaka University, president of the Bangladesh Psychometric Association, and secretary general of the Bangladesh Psychological Association.
Views expressed in this article are the author's own.
Follow The Daily Star Opinion on Facebook for the latest opinions, commentaries, and analyses by experts and professionals. To contribute your article or letter to The Daily Star Opinion, see our guidelines for submission.



Comments