Watson-Glaser Critical Thinking Test 2026: The Complete Guide for Law & Consulting Interviews
If you've got a training contract application, a vacation scheme, or a consulting first-round interview coming up, there's a good chance a recruiter has already mentioned three words that make most candidates quietly panic: Watson-Glaser test. The Watson-Glaser test in 2026 is still the single most common critical-reasoning assessment used by law firms worldwide, and it's spreading further into consulting and other graduate schemes every year — which means understanding exactly what it measures, how it's scored, and how to prepare for it isn't optional anymore. It's table stakes.
This guide covers everything: what the test actually is, the item-banked format and the five sections underneath the RED model, who uses it and why, the percentile-based scoring system (there's no fixed "pass mark," and almost everyone gets that wrong), sample questions with the reasoning behind each answer, and a week-by-week prep plan you can actually follow.
What is the Watson-Glaser Critical Thinking Test?
The Watson-Glaser Critical Thinking Appraisal (WGCTA) is one of the oldest psychometric assessments still in active commercial use. Goodwin Watson and Edwin Glaser, both at Columbia University's Teachers College, first published it as the "Watson Test of Fair-Mindedness" back in 1925, with the first full administration in 1928. It was reworked into the Watson-Glaser Critical Thinking Appraisal in the early 1940s and has been periodically revised since — most recently as Watson-Glaser III, published in 2011 by Pearson TalentLens, which now owns and distributes it.
That 2011 revision is the version almost every employer uses today. It introduced more business-relevant scenarios (rather than the more academic, abstract passages of earlier editions), moved scoring onto Item Response Theory (IRT) rather than simple raw-score norms, refreshed the norm groups the test compares you against, and added an online retest option that publishers can use to validate an in-person result. In other words: this is not a test that was thrown together to keep up with a hiring trend. It has a 90-plus year track record, and its persistence says something about what it's actually trying to measure — not vocabulary, not case-cracking speed, but whether you can reason cleanly from evidence to a conclusion without smuggling in assumptions you didn't notice you were making.
That's also exactly why the test rewards a specific kind of practice. It isn't testing subject knowledge you can cram the night before. It's testing a reasoning habit — and habits are built through repetition under time pressure, not last-minute review.
The format: 40 questions, five sections, one RED model
The modern Watson-Glaser III is item-banked: candidates don't all see identical questions, they draw from a large pool of items calibrated to the same difficulty and construct, which is part of why memorized "answer sheets" circulating online are so unreliable — the specific scenario in front of you may never appear in any leaked set.
The structure itself, though, is fixed and well documented:
- 40 questions total
- Roughly 30 minutes allotted (some employers extend this slightly, but 30 minutes is the standard Pearson administration)
- That works out to about 45 seconds per question — brutally tight once you factor in reading time for the passage each question is based on
Those 40 questions are split across five sections, and the five sections are themselves organized under what Pearson calls the RED model of critical thinking: Recognize assumptions, Evaluate arguments, Draw conclusions. Three of the five sections (Inference, Deduction, and Interpretation) all fall under "Draw conclusions" — they're testing slightly different flavors of the same underlying skill, reasoning from stated facts to a sound conclusion without overreaching.
Here's how the 40 questions break down section by section.
Inference (5 questions)
You're given a short passage of facts to accept as true, followed by a series of proposed conclusions. For each one, you choose from a five-point scale: True, Probably True, Insufficient Data, Probably False, or False. This is the section most people find conceptually hardest, because it's testing your ability to judge the strength of evidence, not just whether a conclusion is possible. "Probably true" and "insufficient data" get confused constantly — more on that in the common-mistakes section below.
Recognition of Assumptions (12 questions)
The largest section alongside Evaluation of Arguments. You're shown a statement, followed by several proposed assumptions underlying it, and for each you decide simply: assumption made, or assumption not made. An assumption is something the statement's argument depends on being true, even though it isn't stated outright. This section punishes over-thinking — candidates often import assumptions from real-world plausibility rather than sticking strictly to what the specific statement requires.
Deduction (5 questions)
Deduction questions give you premises and a proposed conclusion, and you decide whether the conclusion follows or does not follow — purely as a matter of logical necessity, using only the information given. This is the section where you have to switch off common sense entirely. A conclusion can feel true in the real world and still not "follow" from the stated premises, and the test wants you to notice that gap every time.
Interpretation (6 questions)
You're given a paragraph of information and a proposed conclusion, and you judge whether it follows beyond reasonable doubt based on the passage — again, a binary judgment (follows / does not follow), but applying a slightly different standard than Deduction because Interpretation passages are usually descriptive or statistical rather than strict logical premises.
Evaluation of Arguments (12 questions)
The other large section. You're given a question of policy (something like "Should X be done?") followed by a series of arguments for or against, and for each you judge whether it's a strong argument or a weak argument — based on relevance and importance to the question, not on whether you personally agree with it. This is the section that trips up people who bring in their own opinions; the test wants you to evaluate the argument's logical force, independent of your position on the underlying issue.
How the Watson-Glaser test is scored in 2026
This is the part candidates most consistently get wrong: there is no fixed pass mark. Unlike a driving theory test where 43/50 is always a pass, the Watson-Glaser is scored on a percentile basis, comparing your raw score against a relevant norm group (typically graduates or a specific professional population). What counts as "competitive" depends entirely on which employer is using the results and how selective their process is.
That said, published benchmarks give a reasonably consistent picture of where the percentile bands fall:
- Around 33–34 correct out of 40 typically maps to roughly the 80th percentile
- 36–38 correct tends to land around the 90th percentile
- 39–40 correct — near-perfect — usually puts you in the 95th–99th percentile
Most competitive employers — Magic Circle law firms in particular — are looking for candidates at the 75th–80th percentile or higher, and some London firms have historically set an informal bar around a 75% overall correct-answer proxy, though because scoring is norm-referenced rather than raw-score-referenced, the honest takeaway is: aim to answer as many correctly as you can inside the time limit, don't obsess over hitting one specific number, and understand that a "good" score at one firm may just clear the bar at another. For a deeper breakdown of how firms apply these percentiles in practice, AssessmentDay's Clifford Chance profile and The Lawyer Portal's guide to passing the Clifford Chance Watson Glaser test are both useful for firm-specific context.
One more scoring detail worth knowing: because the test is IRT-based and item-banked, your score reflects both how many you got right and the calibrated difficulty of the specific items you were served — which is another reason generic "cheat sheets" with fixed answers are a bad use of your prep time. You're better off building the underlying skill than memorizing outputs to inputs you probably won't see.
Who uses the Watson-Glaser test
Law firms — where it started, and where it's still most concentrated
The Watson-Glaser is overwhelmingly a UK law-firm test. It's the default critical-thinking screen for training contract and vacation scheme applications at most of the Magic Circle and a long tail of other commercial firms, including:
- Clifford Chance
- Linklaters
- Freshfields Bruckhaus Deringer
- Hogan Lovells
- Allen & Overy
- Norton Rose Fulbright
- DLA Piper
- Simmons & Simmons
- Baker & McKenzie
- The UK Government Legal Profession's Legal Trainee Scheme
It typically shows up right after an online application, before assessment centers and interviews — a hard early filter that a large share of candidates never get past because they treat it as an afterthought. The Oxford University Careers Service's guide to preparing for the Watson-Glaser is a genuinely good, non-commercial resource if you want a university-careers-office perspective on how seriously to take it.
Consulting and other graduate schemes — a smaller but growing footprint
Outside law, the Watson-Glaser (or close equivalents built on the same construct) turns up in select consulting recruitment processes and broader graduate schemes, including in India and the US, where the underlying skill — reasoning cleanly from data to a defensible conclusion without overreaching — maps directly onto case-interview thinking. It's far less universal in consulting than SHL or bespoke case-style numerical tests, but candidates applying to graduate programs at professional-services and financial firms increasingly report seeing it, or a Pearson-branded critical-thinking variant of it, somewhere in the process. If you're also prepping for case interviews at MBB firms, see our McKinsey/BCG/Bain case interview guide — the structured-reasoning muscle you build for Watson-Glaser (isolate the claim, isolate the evidence, don't let the two blur) is close to directly transferable to case-interview logic.
Sample Watson-Glaser questions by section, with reasoning
Practicing with real answer logic — not just answer keys — is what actually moves your score. Here's one worked example per section.
Inference example
Passage: "A survey of 500 employees at a mid-sized firm found that 68% reported feeling more productive working from home at least two days per week. The firm subsequently expanded its hybrid-work policy."
Proposed conclusion: "Every employee at the firm prefers working from home."
Correct judgment: False. The passage states 68% reported feeling more productive under hybrid arrangements — it says nothing about "every employee," and in fact directly implies roughly a third did not report this. Where candidates go wrong: they see "productivity" language and a policy change and assume general enthusiasm, without checking whether the specific claim ("every employee") is actually supported. The fix is mechanical: isolate the exact wording of the proposed conclusion and check it word-for-word against the passage, not against what feels plausible.
Recognition of Assumptions example
Statement: "We should hire the candidate with the highest test score, since our assessment process reliably identifies the best performers."
Proposed assumption: "Test scores from the assessment process correlate with actual job performance."
Correct judgment: Assumption made. The statement's recommendation ("hire the highest scorer") only makes sense if the underlying premise — that the assessment reliably predicts performance — is being taken as true. Candidates often mark this "not made" because it feels too obvious to count as an assumption; but obvious, load-bearing premises are exactly what this section is testing you to spot.
Deduction example
Premises: "All partners at the firm have completed a training contract. Sarah has completed a training contract."
Proposed conclusion: "Sarah is a partner at the firm."
Correct judgment: Conclusion does not follow. This is a classic affirming-the-consequent error: all partners have completed a training contract, but that doesn't mean everyone who has completed a training contract is a partner. It's one of the most common logical fallacy shapes the Deduction section tests, and once you recognize the pattern, similarly structured questions get much faster to answer.
Interpretation example
Passage: "In a study of 200 law firms, those with structured mentoring programs reported 15% lower first-year associate attrition than those without. The study did not control for firm size or practice area."
Proposed conclusion: "Structured mentoring programs directly cause lower attrition."
Correct judgment: Does not follow beyond reasonable doubt. The passage describes a correlation and explicitly flags an uncontrolled variable — this is the section's signature move: giving you a plausible causal-sounding conclusion built on data that only supports correlation.
Evaluation of Arguments example
Question of policy: "Should law firms require all trainees to rotate through at least four practice areas?"
Proposed argument: "Yes, because rotation exposes trainees to a broader range of legal work, helping them make a more informed choice of eventual specialism."
Correct judgment: Strong. It's directly relevant to the policy question and addresses a real, important consequence (informed specialization choice). A weak argument on the same question might be something like "No, because some trainees find moving desks annoying" — true perhaps, but trivial and not proportionate to the actual policy decision.
A structured Watson-Glaser prep plan
Because this test rewards a reasoning habit rather than memorized content, cramming the night before rarely works. A two-to-three week structured plan does.
Week 1 — Learn the five question types cold. Don't touch full timed tests yet. Do 10–15 untimed questions per section per day, and for every single one — right or wrong — write one sentence explaining why the correct answer is correct. This is the step almost everyone skips, and it's the one that actually builds the skill rather than just familiarity with the format.
Week 2 — Introduce timing, section by section. Time yourself on Recognition of Assumptions and Evaluation of Arguments first (the two 12-question sections carry the most weight and the most volume), aiming for under 45 seconds per item. Then move to the shorter Inference and Deduction sections, which often need slightly more time per question because of the five-point Inference scale and the strict logic of Deduction.
Week 3 — Full 40-question, 30-minute simulations. Run at least three to five full-length timed tests under realistic conditions — same room you'll test in if you can, no phone, no pausing. After each one, don't just check your score: go back through every wrong answer and categorize the mistake (misread the passage, imported an outside assumption, confused "probably true" with "insufficient data," ran out of time). Patterns in your errors matter more than the raw score itself at this stage.
The final 48 hours — light review only. Re-read your error log from Week 3, skim the RED model breakdown one more time so the five section rules are fresh, and stop doing new full-length tests — you want to walk in reasoning clearly, not fatigued from over-practice.
If you want a broader view of how this fits into your overall interview and assessment prep — building STAR stories for the interview stages that typically follow a Watson-Glaser screen, or mapping out the rest of a firm's process — ClavePrep's STAR story builder and our full suite of prep tools are built for exactly that kind of end-to-end preparation, and how ClavePrep works walks through the full method.
Common mistakes candidates make
Confusing "probably true" with "insufficient data" in Inference. These two options are the most frequently mixed up on the entire test. "Probably true" means the evidence leans that way but doesn't fully confirm it; "insufficient data" means you genuinely can't tell either direction from what's given. If more information would change your confidence meaningfully in one direction, it's "probably true" or "probably false," not "insufficient data."
Bringing outside knowledge into Deduction. Deduction questions want strict logical validity, full stop. If a conclusion is true in the real world but doesn't logically follow from the specific premises given, the correct answer is still "does not follow." Real-world truth and logical validity are not the same test.
Letting personal opinion bleed into Evaluation of Arguments. The test isn't asking whether you agree with an argument's position — it's asking whether the argument, as stated, is logically strong or weak in relation to the policy question. Strongly-held personal views are the single biggest distortion source in this section.
Under-practicing Recognition of Assumptions. It's 12 of 40 questions — 30% of your total score — and it's also the section most candidates skip in favor of the "flashier" Inference and Deduction sections. Weight your practice time to match the test's actual weighting.
Treating leaked "answer sheets" as real prep. Because the test is item-banked, a memorized answer key for one candidate's test rarely matches yours. Time spent memorizing answers is better spent internalizing the reasoning rules above.
Ignoring the clock until test day. At roughly 45 seconds per question with reading time built in, pacing has to become automatic well before the real assessment. Untimed accuracy and timed accuracy are genuinely different skills, and only one of them is being measured.
Frequently asked questions
Is the Watson-Glaser test hard to pass? There's no single pass/fail line — it's scored on a percentile basis against a norm group, so "hard" depends on which employer's bar you're aiming for. Competitive law firms typically want candidates at the 75th–80th percentile or above, which in practice usually means getting somewhere around 33+ of the 40 questions correct, though this varies by norm group and firm.
How many questions are on the Watson-Glaser test, and how long do I get? 40 questions across five sections, with roughly 30 minutes allotted in total — about 45 seconds per question on average, though some sections (like the five-point Inference scale) can eat more time per item than others.
What is the RED model in the Watson-Glaser test? RED stands for Recognize assumptions, Evaluate arguments, and Draw conclusions — the three underlying critical-thinking constructs Pearson uses to organize the five scored sections. Recognition of Assumptions and Evaluation of Arguments map directly to "Recognize" and "Evaluate"; Inference, Deduction, and Interpretation all sit under "Draw conclusions."
Which law firms use the Watson-Glaser test? Most Magic Circle and major UK commercial firms, including Clifford Chance, Linklaters, Freshfields Bruckhaus Deringer, Hogan Lovells, Allen & Overy, Norton Rose Fulbright, DLA Piper, and Simmons & Simmons, along with the UK Government Legal Profession's trainee scheme. It's less universal outside the UK but does appear in some international consulting and graduate-scheme processes, including in India and the US.
Can I retake the Watson-Glaser test if I fail? Policies vary by employer. Some firms allow a retake after a set cooling-off period (often six to twelve months), especially if you're reapplying for a future recruitment cycle; others don't. Check the specific firm's graduate recruitment FAQ before assuming either way.
What's the difference between Watson-Glaser Inference and Deduction questions? Inference questions allow you to weigh evidence and probability — you're judging how strongly a conclusion is supported on a five-point scale from true to false. Deduction questions are strictly binary and logic-only: does the conclusion follow from the stated premises as a matter of necessity, with no room for probabilistic judgment or outside common-sense knowledge.
Is the Watson-Glaser test the same as an IQ test? No. It measures a specific, trainable skill set — recognizing assumptions, evaluating argument strength, and drawing sound conclusions from stated evidence — rather than general cognitive ability. That's also good news for prep: because it's a defined, practiced skill rather than a fixed trait, structured practice measurably moves your score.
How is the Watson-Glaser different from other reasoning tests like SHL or CCAT? Where general critical reasoning or verbal-numerical batteries like SHL test a broader mix of skills, the Watson-Glaser is narrowly focused on argument and inference logic, using the specific five-section, RED-model structure described above. It's also far more concentrated in one industry — law — than most other major assessments, which is why firm-specific prep (knowing exactly which firms use it and how they weight it) matters more here than for broader aptitude tests.
Getting ready for test day
The Watson-Glaser rewards exactly the kind of preparation most candidates skip: slow, deliberate practice on each of the five question types before you ever add a timer, followed by realistic full-length simulations once the underlying logic is second nature. Treat it as a skill you're building, not a trivia set you're memorizing, and the percentile climb takes care of itself.
If you're building out a fuller prep plan — for the Watson-Glaser itself, the interviews that typically follow it, or a parallel consulting application — ClavePrep's interview prep tools and STAR story builder are designed to carry you from assessment through to offer, and how ClavePrep works has the full picture of the method behind it.
