The method

Everyone sells more questions. Almost nobody builds the part that decides whether they stick.

Question count is the easiest thing to photograph and the least important variable in the research. What moves retention is when you practise and what the practice asks of you — findings that have been settled for decades and are almost never implemented, because they make studying feel worse while making performance better. This page is the list, the papers behind it, and an honest column marking what is in the product today.

Why prep is built the way it is

Volume sells because it is legible. Five thousand questions is a number a parent can compare across two websites in four seconds. The scheduling of those questions is not comparable, is not visible on a sales page, and is the part the evidence is about.

There is a second reason, and it is less flattering. The methods that work produce worse practice sessions. You get more wrong. It feels slower. Re-reading an explanation until it is familiar feels like progress and produces very little; being asked the same thing four days later and failing to recall it feels like going backwards and is the thing that builds the memory. Robert Bjork named this class of intervention desirable difficulties, and the word doing the work is desirable.

A product that optimises for how studying feels will build the wrong thing every time. So the rest of this page is what we built instead, what it came from, and where we have not got to yet.

Seven decisions, and where each came from

Each one is labelled with what it does today: four in your hands right now, one built and waiting on a surface, two not built at all. We would rather be dull about status than have you discover the gap yourself.

  • In the product today
  • Built, not yet in your hands
  • Not built yet
01

Answering beats re-reading

In the product today

The finding

Students who read a passage and then took recall tests on it, versus students who read it three more times, showed the pattern that defines this whole field: the re-readers did better after five minutes and much worse a week later. Producing an answer from memory is what builds the memory. Reading the explanation again mostly builds the feeling of knowing it.

Roediger & Karpicke (2006), Psychological Science

What we do with it

There is no passive mode here. Every academic surface asks you to produce an answer before it shows you anything — 2,156 questions across four banks, sat as a full-length, a twenty-five question focused set or a twelve-question warm-up, with the review built from the questions you actually sat rather than the whole bank. That is the cheap part of the finding, and every prep product gets it right by accident. The expensive part is the next four.

02

The gap between reps is the variable nobody sells

Built, not yet in your hands

The finding

Across hundreds of experiments, the same total study time spread over days beats the same time massed into one sitting, and the effect is not small. Later work in the same line found the best gap scales with how far away the test is — the further out the exam, the wider the gaps should be.

Cepeda, Pashler, Vul, Wixted & Rohrer (2006), Psychological Bulletin

What we do with it

Every question you answer already gets a return date on an expanding ladder: 1, 2, 4, 8, 16, 30 days. Get it right while sure and it moves up a rung. Miss it and it drops to the bottom. That much is running now — the schedule is being written from your answers as you sit tests. What is not in your hands yet is the session that hands those items back to you: it is built, and it is not yet linked anywhere you can reach. Until it is, this is a queue accumulating in the background, and we are not going to pretend otherwise. The whole ladder is spelled out below — we would rather show it than describe it.

03

One tap before the reveal

In the product today

The finding

Students are poor judges of what they know, and the errors run in a specific direction: fluent material feels learned when it is not. Asking for a judgement BEFORE feedback is what makes the miscalibration visible — to the student and to anyone helping them. There is a second finding underneath it, and it is the surprising one: errors made with high confidence are the ones most likely to be corrected once the right answer arrives.

Dunlosky & Rawson on calibration; Butterfield & Metcalfe (2001) on hypercorrection

What we do with it

Before the answer is revealed, one tap: sure, think so, or guessing. It is skippable, it is never scored, and it costs one interaction to split four genuinely different events that a score report records identically. The one it exists for is sure-and-wrong — a misconception you are certain of. That is the highest-value thing in any tutoring session and the single thing a percentage score can never surface, because a confident error and a coin-flip error are the same red X.

04

Three ways to miss a question, three different fixes

In the product today

The finding

This one is ours, not a citation. There is no study that says careless, concept and timing are the right three buckets. What the literature does support is the general shape: a remedy has to match the cause, and undifferentiated practice on everything you got wrong spends most of its time on the wrong problem.

No citation — this is a design decision, and we are labelling it as one

What we do with it

After a miss, one question: what happened? Careless, didn't know it, ran out of time — plain words on the screen, three different remedies behind them, and dismissible, because an honest "I'd rather not say" is worth more than a shrug recorded as data. Careless is a checking and pacing problem and more content teaching will not touch it. Concept needs instruction before more reps. Timing means the section strategy is wrong — the question was answerable and you never got a fair look at it. Same red X, three different weeks of work.

05

Mixing problem types — partly done, and we will say which part

Not built yet

The finding

Practising one problem type until it is comfortable, then moving to the next, produces better practice sessions and worse test performance than shuffling the types together. Blocked practice lets you recognise the method from the position on the page; mixed practice forces you to choose the method, which is the thing the test actually asks.

Rohrer & Taylor on interleaved mathematics practice

What we do with it

What ships today: within a section, questions come off a shuffled pool, so you get skills mixed rather than five geometry items in a row. What does NOT ship today: deliberate interleaving. Nothing in the selector is choosing to alternate skills, and a sitting is still blocked by section because a real SAT is. The review queue will improve on that almost by accident — it orders by urgency rather than by section, so a Reading item lands next to two Math ones — but ordering something incidentally is not designing for it. Building practice by design, this skill then a different one then back, is the next piece of work, and until it exists we are not going to call a shuffle an interleaving engine.

06

Practice that knows you had two-a-days

In the product today

The finding

No study to cite, and the honest reason is that the studies were not run on athletes in season. What the spacing literature does establish is that the schedule is the active ingredient — which means a schedule built for someone with free evenings is not a small mismatch for an athlete with a 6am lift and a Friday bus. It is the wrong intervention.

Extension of the distributed-practice work above, not a separate finding

What we do with it

Your plan comes in two shapes and you choose which one you are in: in-season, built from short sessions that survive a practice schedule, and off-season, built for a real workload. Same six weeks, same target section, different geometry. It is a toggle you flip, not something we infer — driving the switch automatically from your team calendar is where this goes next, and so is the target test date the scheduler is already written to respect.

07

Ten minutes of writing before the test

Not built yet

The finding

In a classroom experiment, students who spent ten minutes writing about their worries immediately before a high-stakes exam outperformed a control group, and the benefit was concentrated in the students who were most anxious about testing. It is one of the few interventions in this whole literature that costs ten minutes and nothing else.

Ramirez & Beilock (2011), Science

What we do with it

Not built yet. Saying so plainly is the point of this page. It belongs here more than it belongs anywhere else — an athlete already understands a pre-game routine, and this is the same object with a pen. When it ships it will be ten minutes, unscored, unshared, and it will not go anywhere near your profile.

The schedule, in full

The numbers below are the actual constants in the scheduler, not an illustration of them. If you are going to trust a queue to decide what you work on, you should be able to read the rule it uses. Worth repeating from 02 before you read it: this rule is running on your answers today, and the session that hands the items back is the piece still to land.

The ladder

Answer something correctly and sure, and it climbs. Every rung roughly doubles the gap, which is the shape the distributed-practice work points at: as a memory holds, the useful test of it moves further away.

  1. 1dayRung 1
  2. 2daysRung 2
  3. 4daysRung 3
  4. 8daysRung 4
  5. 16daysRung 5
  6. 30daysRung 6

A miss resets it. Back to tomorrow, whatever the streak was. Partial credit for a streak that just broke would be scheduling around a memory you demonstrably do not have.

It is written to stop at your test date — if the next interval would land after the exam, the item leaves the queue, because a queue full of items due next month reads as noise on the night before. Being exact: there is nowhere in the product to tell us your test date yet, so that rule is written and tested and currently does nothing.

Thirty days is the ceiling. A prep season is weeks, and an algorithm that tunes a per-item difficulty over dozens of reviews would be fitting noise on a horizon this short.

What one tap changes

A right answer and a right answer are not the same event. Every row below is a different thing to do next, and a score report collapses all five into two.

You tappedResultComes backWhy
SureCorrectUp one rungReal knowledge. Push it out — 2 days, then 4, then 8, 16, 30.
ThinkCorrectUp one rungTreated exactly like sure. Hedging on a right answer is ordinary calibration, not a gap.
GuessCorrect2 daysLuck, not learning. A one-in-four chance should not buy the same eight-day gap as knowing it.
GuessWrongTomorrowAn ordinary gap. Short, but not urgent — you already know you do not know it.
SureWrongTomorrow, at the frontA misconception you are certain of. It jumps the queue whatever your streak was, because being wrong while certain means you have no reason to look it up.

Two design notes we will hold ourselves to. Every answer is written to an append-only log, and the queue is derived from it — so when the algorithm improves, and it will, the schedule is rebuilt from your real history rather than migrated in place. And the tap is never scored. Admitting a guess has to be free, or the honest answer stops being the easy one and the whole mechanism is worth nothing.

Three ways to miss a question

The percentage tells you how many. It cannot tell you which kind, and the kind is what decides the next six weeks. This is the one place on the page where we are following a design judgement rather than a study, and we would rather flag that than borrow a citation that does not fit.

Careless

You knew it and lost it anyway.

A checking and pacing problem. More instruction on the concept is the wrong prescription and the most commonly prescribed one — you already have the concept. What changes this is a checking habit and a slower final look, not another lesson.

Concept

You did not know the underlying idea.

The only one of the three where more teaching is the answer. This is where a worked example earns its place — walking through a solved problem beats attacking a blank one when you are new to a topic, and stops paying once you are not.

Timing

You could have answered it, with time.

A section strategy problem, and it is invisible in a score. Ten questions you never reached and ten questions you got wrong produce the same number. The fix is order and triage, not content.

What we don't do

A method page that only lists what a company believes in is an advertisement. These are the four things we have decided not to build or not to say, and each one costs us something.

No learning styles

There is no visual-learner setting here and there never will be. The idea that matching instruction to a preferred style improves learning is one of the most tested claims in education, and studies designed properly to test it — the design where the same material is taught both ways to both groups — do not find the effect. People do have preferences. Teaching to them does not help. Any product that asks whether you are a visual or auditory learner has told you it does not read the research.

Pashler, McDaniel, Rohrer & Bjork (2008), Psychological Science in the Public Interest

No score-gain guarantee. Not now, not later.

We have never run a controlled trial on our own users, so we do not know what PlayEligible does to a score, and neither does anyone quoting you a number. A promise of plus-150 points is a marketing figure measured on somebody else, if it was measured at all. We cite the research and we build from it; we do not convert another team's effect size into a promise about you. If we ever run a study of our own, the design goes up before the results do.

We do not dress up the study habits that do not work

The big review of ten common study techniques rated only two as high-utility across the board: practice testing and distributed practice. Highlighting, re-reading and summarising rated low. That is inconvenient, because highlighting is pleasant and easy to build, and testing yourself is neither. We built the two that work and left the rest out.

Dunlosky, Rawson, Marsh, Nathan & Willingham (2013), Psychological Science in the Public Interest

No motivational layer sold as science

Growth mindset is real and it is oversold. The largest national trial found roughly a tenth of a grade point, concentrated in lower-achieving students, and a 2018 meta-analysis found average effects that are small. That is a genuine result and it is not a product. We are not going to put a mindset module in front of you and call it evidence-based.

Yeager et al. (2019), Nature; Sisk et al. (2018), Psychological Science

Sources

Every one of these is a real, published paper you can look up. Where we could not place a citation with confidence we described the finding and left the reference off, and where a decision is ours rather than the literature's we said so above. If you think we have misread one of these, tell us and we will fix the page.

  1. Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science.

    Retrieval beats restudy at a delay, and loses to it immediately — which is why the wrong method feels better.

  2. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin.

    The spacing effect, across a very large body of experiments. The basis for the ladder.

  3. Rohrer, D., & Taylor, K. Work on interleaved and distributed practice of mathematics problems.

    Mixing problem types depresses practice performance and improves delayed test performance.

  4. Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition.

    Desirable difficulties: the conditions that slow acquisition and improve retention. The frame for everything above.

  5. Butterfield, B., & Metcalfe, J. (2001). Errors committed with high confidence are hypercorrected. Journal of Experimental Psychology: Learning, Memory, and Cognition.

    Confident errors are corrected more readily once feedback arrives — the reason the sure-and-wrong tap is worth its interaction.

  6. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest.

    Ten techniques rated. Practice testing and distributed practice are the two that come out high-utility.

  7. Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction.

    Worked examples beat problem solving for novices — the finding behind leading with a solved problem on a concept miss. Later work found the advantage shrinks and can reverse as expertise grows, which is why explanations should thin out rather than scale up.

  8. Ramirez, G., & Beilock, S. L. (2011). Writing about testing worries boosts exam performance in the classroom. Science.

    Ten minutes of expressive writing before an exam, with the benefit concentrated in anxious students.

  9. Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008). Learning styles: Concepts and evidence. Psychological Science in the Public Interest.

    The meshing hypothesis does not survive studies designed to test it.

  10. Yeager, D. S., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature.

    A real but small effect, concentrated in lower-achieving students. Cited here so nobody has to take our word for the size.

  11. Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L., & Macnamara, B. N. (2018). To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science.

    Average growth-mindset effects across the literature are small, with the clearest benefits for students at academic risk.

One more time, because it is the whole point: none of this is a claim about what will happen to your score. These are findings from other people's controlled studies, and we have built the product they point at. We have not measured ourselves, and we are not going to sell you a number we did not measure.

Start with one timed sitting.

Everything above needs a first rep to schedule. Fifteen minutes is enough to give the queue something real to work with, and a full-length is free.

The method — what actually works in test prep | PlayEligible