Dual N-Back
English
← Back to the app

Dual N-Back · Guides

How dual n-back scoring works: hits, misses and false alarms

Why a single percentage hides most of what happened in a session, how the four possible outcomes work, and which numbers are worth writing down.

Dual N-Back · Memory & attention
Dual N-Back · Memory & attention

Two decisions on every turn

Each turn of a dual n-back presents a position and a sound at the same time, and asks two independent questions. Is this position the same as the one N turns back? Is this sound the same as the one N turns back? You answer each one separately, usually with a different key or button, and crucially you also answer by doing nothing. Staying still is a real answer, not an absence of one. Most explanations of the task skip this, and it causes a lot of confusion for beginners, who assume they are only supposed to act when something matches. In signal-detection terms, every turn is two yes-or-no judgements under uncertainty, and how you handle the uncertain middle is what your score is really measuring.

The four outcomes

Because each judgement is a yes or no, and because the truth is also a yes or no, there are exactly four things that can happen. A hit is a match you correctly identified. A miss is a match you failed to press for. A false alarm is a press when there was no match. A correct rejection is staying still when there was nothing there. Hits and correct rejections are good, misses and false alarms are errors, and the two kinds of error are not interchangeable. Missing a match usually means the item slipped out of your window. Pressing when nothing was there usually means you guessed, or that a similar item from a different distance felt close enough. They point at different problems and they call for different corrections.

Why a single percentage misleads

Many implementations show one number, often something like percentage correct. The trouble is that correct rejections are by far the most common outcome. In a typical block, matches are deliberately rare, so most turns contain nothing in either stream. If you simply never pressed at all, you would still score highly on a naive percentage, because you would collect a correct rejection on almost every turn. That is why a raw percentage can look respectable after a session in which you barely engaged, and why comparing two sessions on that number alone can tell you the opposite of what happened.

The bias problem

Every player sits somewhere on a scale between pressing too rarely and pressing too often. Press rarely and your false alarms fall while your misses rise. Press often and the reverse happens. Neither end is skill: both are a choice about how much uncertainty you are willing to act on, and you can move along that scale without your memory changing at all. This is the single most important thing to understand about your own scores. If your level went up this week, the first question to ask is whether your accuracy improved or whether you simply became more willing to press.

Reading your two streams separately

The position stream and the sound stream fail differently, and averaging them hides that. Many people are noticeably stronger on position, because a lit square on a grid has a spatial anchor that a spoken letter does not. Others find the reverse. If your app reports each stream, look at them apart before looking at the total. A session where position is clean and sound is chaotic is a different problem from one where both are equally shaky, and the first calls for time spent on the sound stream alone rather than more general practice.

What a good session actually looks like

Rather than a target percentage, which depends on the implementation, use a shape. A good session has a high hit rate, a low false-alarm rate, and roughly similar performance in the first and second halves. A session that starts strong and falls apart is too long. A session with a high hit rate and a high false-alarm rate is a guessing session, whatever the headline number says. A session with very few of both is one where you were too cautious to learn anything, which is less common but does happen when people overcorrect after a bad week.

Session-to-session noise

This task is unusually sensitive to condition. Sleep, caffeine, time of day, background noise and how recently you ate all move scores by an amount that can easily exceed a week of genuine improvement. That means a single session is close to meaningless as a measurement, and that comparing your Tuesday morning result with your Friday midnight result tells you about your Friday midnight, not your memory. The practical consequence is simple: look at groups of sessions, and try to keep the conditions roughly constant when you can.

What is worth writing down

Four things are enough: the level, the hit rate, the false-alarm rate and a one-word note on your state, such as tired, rushed or fine. That last one costs nothing and explains most of the outliers you will otherwise puzzle over. If you want a fifth, record the pace, because a change there invalidates comparison with everything before it. What is not worth recording is your best-ever level, which is a trophy rather than a measurement.

Comparing yourself with yourself

There is no standard dual n-back, no agreed pace, no agreed number of trials and no verification anywhere. A level reported online could have been produced with a slow interval, a single stream and generous guessing. Because of that, the only comparison with any meaning is between your own sessions run under settings you have not changed. This is less exciting than a leaderboard and it is the only version of progress in this exercise that can be trusted.

Reading a plateau through the numbers

When your level stops moving, the numbers usually explain why. If hits are flat and false alarms are flat, you are genuinely at your ceiling for now, and the answer is time rather than a new method. If hits are flat while false alarms creep up, you are compensating with guesses and the level is at risk of drifting above your real ability. If both are falling, look at your sessions rather than your memory: length, timing and sleep account for most of it.

Using the numbers without becoming a spreadsheet

There is a point past which measuring the exercise replaces doing it. Four numbers and a word take fifteen seconds and give you everything you need. Anything more elaborate tends to be a way of feeling productive without practising. Dual N-Back shows your session results and lets you review your practice history, which is enough to see whether your accuracy is moving in the direction you want, at a level that genuinely stretches you.

Questions

What is a false alarm in n-back?

Pressing when there was no match. It is the error most people never look at, and it is where guessing hides. A session with fewer hits and far fewer false alarms is usually better than one with more of both.

What percentage is good in dual n-back?

No fixed percentage means much, because implementations differ in how they count. A better test is the shape: high hits, low false alarms, and similar performance in the first and second half of the session.

Why does my score change so much day to day?

This task is unusually sensitive to sleep, caffeine, time of day and background noise, and those move scores by more than a week of genuine improvement. Compare groups of sessions run under similar conditions, not single results.

What should I record after a session?

Four things: the level, the hit rate, the false-alarm rate and one word about your state. If you change the pace, record that too, because it invalidates comparison with everything before it.

Further reading