Dual N-Back · Guides
Reaction time and the Stroop test: what these tasks actually measure
How a reaction-time task differs from a Stroop task, what each one loads, why your scores swing so much, and how to compare results without fooling yourself.

Two tasks that feel alike and are not
Reaction time and Stroop are often grouped together as quick brain games, and they do share a surface: both are short, both give you a number in milliseconds, and both feel mentally effortful. Underneath they measure different things. One measures how fast you can respond at all. The other measures what it costs you to produce the right answer when an easier wrong answer is already available. Confusing them is the most common reason people misread their own results.
A reaction task starts with a signal
In a simple reaction-time task, you wait for a signal and respond as fast as you can. There is nothing to decide and nothing to remember. What you are measuring is the whole chain from the signal reaching your eye to your finger moving, which includes perception, the decision that something happened, and the motor response itself. It is the cleanest measure in this family precisely because it removes choice.
Choice reaction time adds a decision
Add options and the task changes. If two signals can appear and each has its own response, you are now measuring decision time on top of everything else, and your times lengthen. The more alternatives there are, the longer the average response, which is one of the oldest regularities in experimental psychology. This is why a choice task and a simple task are not comparable, even on the same device on the same day.
What reaction time is not
It is not a measure of intelligence, of skill, or of how sharp you are as a person. It is a measure of a particular chain of events on a particular device at a particular moment. A large part of any score is the hardware and the software: touch screens, refresh rates and input polling all add delay that has nothing to do with you, and that delay differs between phones.
Why your reaction times swing so much
Reaction time is among the most state-sensitive measures you can take. Sleep, caffeine, time of day, temperature, how recently you ate and whether you are sitting comfortably all move it, and they can move it by more than any amount of practice will. That makes it excellent as a daily check on your own alertness, and poor as a measure of anything stable over weeks.
The Stroop task, and where it comes from
In a Stroop task you name the colour of the ink a word is printed in, while the word itself spells a different colour. The word RED printed in blue should be answered blue. The task is named after John Ridley Stroop, whose 1935 paper reported that this interference slows people down reliably. It is one of the most reproduced findings in psychology and has been studied for the better part of a century.
Why reading gets in the way
For a literate adult, reading is close to automatic: you do not decide to read a word, it simply happens. Naming a colour is not automatic in the same way. So on an incongruent trial two answers arrive, the automatic one and the correct one, and you must suppress the first to produce the second. That suppression takes measurable time, and the amount of time it takes is what the task is designed to capture.
The Stroop effect, measured
The difference between your times on congruent trials, where the word and the ink match, and incongruent trials, where they conflict, is the Stroop effect. That difference is the number that matters, not your raw speed. MacLeod reviewed half a century of research on this effect in 1991, and the review remains the standard reference for how robust the phenomenon is and how many variations of it exist.
What Stroop measures
Inhibition, or more precisely the cost of producing a controlled response against an automatic one. It is not a memory task. Nothing has to be held or updated, and the correct answer is in front of you the whole time. This is why practising Stroop gives you little reason to expect a change in your memory, and why it sits in a different category from n-back despite both feeling demanding.
How both differ from dual n-back
Dual n-back loads maintenance and updating: you hold a moving window and rewrite it constantly. Reaction time loads speed. Stroop loads inhibition. These are related but separate abilities, and being strong on one predicts surprisingly little about the others. Jaeggi and colleagues examined how n-back performance relates to reasoning measures in 2010, and even there the relationship is more modest and more complicated than popular accounts suggest.
Compare like with like
Almost every comparison people make with these tasks is invalid. Different apps use different timing methods, different trial counts and different rules about what counts as an error. Different phones add different input delay. And your own state changes more between morning and evening than between one month of practice and the next. The only comparison worth anything is your own results, on the same device, under similar conditions.
Reaction time and intelligence
There is a real, well-documented correlation between reaction-time measures and cognitive test scores at population level, examined in a large study by Deary, Der and Ford in 2001. That correlation says something about groups and almost nothing useful about you. A single person's score on a phone, taken once, tells you about that morning rather than about their mind, and no responsible app should present it otherwise.
What to do about errors
Speed alone rewards guessing, because you can always be faster if you stop caring whether you are right. Any honest use of these tasks reads speed and accuracy together, and treats a fast run with several errors as worse than a slightly slower run that was clean. If a task reports only your time, be sceptical of what the number is worth.
Using the numbers as feedback
The useful role for these tasks is as a check on your own state before doing something that matters, and as a slowly moving personal baseline. Take them at roughly the same time of day, on the same device, and look at a week rather than a session. If your reaction time is unusually slow this morning, that is information about this morning, and it is worth having before you judge a dual n-back session done ten minutes later.
Where they fit in Dual N-Back
Dual N-Back includes short cognitive games alongside the main exercise, including reaction time, Stroop and visual memory, and reports your results for each. They are there as practice and as a way to see your own performance, not as an assessment. The app states plainly that scores describe performance on these tasks and that it does not measure IQ or promise an increase in intelligence.
Questions
What does the Stroop test measure?
Inhibition: the cost of producing a controlled answer when an automatic one is already available. Naming the ink colour of a word that spells a different colour forces you to suppress reading. The number that matters is the difference between congruent and incongruent trials, not your raw speed.
What is a good reaction time?
There is no meaningful universal figure, because a large part of any score comes from the device: touch screens, refresh rates and input polling add delay that differs between phones. Compare your own results on the same device under similar conditions instead.
Why does my reaction time change so much?
It is one of the most state-sensitive measures there is. Sleep, caffeine, time of day, temperature and how recently you ate all move it, often by more than practice ever will. That makes it a good daily check on alertness and a poor measure of anything stable.
Does reaction time measure intelligence?
A correlation between reaction-time measures and cognitive test scores exists at population level, examined by Deary, Der and Ford in 2001. It says something about groups and almost nothing useful about one person's score taken once on a phone.
Is Stroop practice the same as memory training?
No. Stroop loads inhibition and nothing has to be held or updated, since the correct answer is visible throughout. Dual n-back loads maintenance and updating. Both feel effortful, which is why they are often confused, but practice on one gives little reason to expect gains on the other.
Further reading
- Stroop (1935), Studies of interference in serial verbal reactions, Journal of Experimental Psychology
- MacLeod (1991), Half a century of research on the Stroop effect: an integrative review, Psychological Bulletin
- Deary, Der & Ford (2001), Reaction times and intelligence differences: a population-based cohort study, Intelligence
- Jaeggi et al. (2010), The relationship between n-back performance and matrix reasoning, Intelligence