The Implicit Association Test (IAT) measures how quickly people sort stimuli when two concepts share a response key. Greenwald, McGhee and Schwartz (1998) introduced it with a simple logic: if "flower" and "pleasant" are strongly associated for a person, sorting flowers and pleasant words with the same key should be faster than sorting flowers and unpleasant words with the same key. The difference in speed between the two pairings is the measure.
Because the measure is a reaction-time difference, an IAT has more procedural requirements than a questionnaire. The block sequence is fixed, trial order is constrained, errors are handled in a particular way, and the order of the two critical pairings has to be counterbalanced across participants. Each of these has to be implemented correctly before the task is moved into a browser. This post goes through those requirements, shows how each one can be built in lab.js, and uses a recent applied study to illustrate the choices involved.
The standard procedure and its common variants
The standard IAT has seven blocks. Greenwald et al. (2022), a best-practice paper written by 22 IAT researchers, gives the typical trial counts as 20, 20, 20, 40, 30, 20 and 40, for 190 trials in total:
- Target practice: sort the two target categories (for example, flowers and insects).
- Attribute practice: sort the two attribute categories (pleasant and unpleasant).
- First combined block, practice: targets and attributes share keys in one pairing.
- First combined block, test.
- Target practice with the target keys reversed.
- Second combined block, practice, with the reversed pairing.
- Second combined block, test.
Block 5 has 30 trials rather than 20 in the current recommendation. IAT scores are known to depend partly on which pairing a participant performs first, and Nosek, Greenwald and Banaji (2005) found that this block-order effect was sharply reduced by adding practice trials. The longer block 5 gives participants that extra practice before the reversed pairing.
Two variants appear often in applied work. The Brief IAT (Sriram & Greenwald, 2009) uses two combined blocks with about a third of the trials, and asks participants to focus on only two of the four categories in each block. The Single-Category IAT (Karpinski & Steinman, 2006) measures associations with one target concept that has no natural opposite category. The lab.js implementation principles below apply to all three; the block counts and category assignments change.
An example from applied health research
A recent example of an IAT-style task built in lab.js comes from nursing home care. Knippenberg, Leontjevas, Declercq, van Lankveld and Gerritsen (2026) asked whether information about a resident's depression changes how caregivers implicitly evaluate, and feel motivated towards, behaviours that could improve the resident's mood. Dutch nursing home caregivers completed two IATs at three time points one month apart (N = 119, 85 and 81). Before the task, they read one of three short vignettes about a resident: no symptoms and no diagnosis, symptoms without a diagnosis, or symptoms with a diagnosis of depression. Information about symptoms predicted implicit motivation (estimate 0.23, 95% CI [0.01, 0.45]), a diagnosis did not add to this, and neither kind of information predicted implicit attitude. The authors ask for cautious interpretation of the motivation effect.
The tasks themselves are described in a companion development paper by the same group (Knippenberg et al., 2025), which reports several design decisions that are relevant to anyone building a similar measure.
- Task variant. The target concept, caregiver behaviours that improve resident mood, has no natural counterpart category. The authors therefore used a single-category variant, with a target category labelled "do something" and attribute categories of positive and negative images (for attitude) or "I want" and "I do not want" images (for motivation).
- Block structure. Five blocks: 16 attribute practice trials, then 32 practice and 64 test trials for one pairing, and 32 practice and 64 test trials for the reversed pairing. The vignette primes were shown between blocks. Because the target category appears with only one attribute at a time, the unpaired attribute category was shown twice as often to balance left and right key presses.
- Responses and errors. Responses were given with the E and I keys. After an incorrect response, a red X appeared in the centre of the screen until the correct key was pressed.
- Counterbalancing. For half the participants, the two pairings were run in the opposite order. The order of the two IATs (attitude first or motivation first) was also randomised. The authors report a block-order effect for the attitude measure (F(1, 223) = 4.56, p = .034).
- Setting. Participation was online and unsupervised. Participants used their own laptop or desktop computer with a keyboard, in a place of their choosing that they judged to be minimally distracting. Nursing homes forwarded an email invitation to their staff.
- Software. The reaction-time tasks were built in lab.js and run on the O4U research platform, with questionnaires in LimeSurvey. The study was not hosted on Open Lab.
- Scoring and reliability. D scores were computed with the built-in error penalty procedure described below, trials over 10,000 ms were removed (0.09% of trials), and one participant was excluded for responding faster than 300 ms on more than 10% of trials. Split-half reliability was .81 and .85 for the two IATs. Test-retest reliability was considerably lower (ICCs of .29 and .25), which is lower than the average of about .50 that Greenwald et al. (2022) report across IAT studies.
Building the block structure in lab.js
In lab.js, an IAT maps onto the builder's standard components. Each block is a Loop whose rows list the stimuli for that block, with a column for the stimulus, its category and the correct key. Inside each loop is a Sequence that forms one trial: typically a short blank interval, the stimulus screen, and the error screen described in the next section. The instruction screen for each block, which shows the category labels and their assigned keys, sits before its loop. The category labels usually stay visible in the upper left and upper right corners of the stimulus screen throughout the block.
On the stimulus screen, the Responses table maps the two keys to response labels (for example,
e to left and i to right), and the correct response field refers to the loop column
that holds the correct key, such as ${ parameters.corr }. lab.js then records the response, the
reaction time and whether the response was correct in each trial's data row. A public lab.js IAT
from the Learning and Implicit Processes Lab at Ghent University
(LIPLabGhent/Past-Nonsuicidal-self-injury-IAT)
follows this layout, with one loop per block and a trial sequence of blank interval, stimulus,
error screen and blank interval.
Greenwald et al. (2022) recommend a brief interval between a response and the next stimulus; 250 ms is common. Knippenberg et al. (2025) used 150 ms after each correct response. In lab.js this is a screen with a timeout and no responses.
Trial order: alternation and runs of the same key
Two trial-order recommendations in Greenwald et al. (2022) need more than the loop's default shuffle. In the combined blocks, target and attribute trials should strictly alternate, and runs of more than four consecutive trials requiring the same key should be avoided.
A loop set to shuffle its rows (the draw-shuffle sampling mode) produces a random order, and a
random order will sometimes produce long same-key runs and will not alternate target and attribute
trials. There are two ways to handle this.
- Pre-generated orders. Generate one or more trial orders that satisfy the constraints in R or Python, and paste each as the loop's rows with sequential sampling. This is transparent and easy to report, but every participant who receives a given list sees the same order.
- Generating the order in a script. A lab.js script that runs before the loop is prepared can build the trial list and assign it to the loop's rows. A public lab.js Brief IAT from the University of Wisconsin–Madison (uwmadison-chm/labjs-biat) uses this approach: it builds balanced lists for the target and attribute categories, shuffles each with lab.js's constrained shuffle helper so that the same item does not appear twice in a row, and then interleaves the two lists so that target and attribute trials alternate. A check on the maximum run length of the same key can be added to such a script.
Error feedback that requires a correction
There are two common ways to handle incorrect responses, and the choice affects scoring. In the first, an error is followed by a red X, and the participant has to press the correct key before the trial ends. The recorded latency then runs from stimulus onset to the correct response, so an error already costs time and no further penalty is added at scoring. Greenwald et al. (2022) describe this built-in error penalty as the preferred procedure. In the second, the X is shown briefly and the trial moves on, and at scoring each error latency is replaced by the block mean of correct trials plus a fixed penalty, usually 600 ms. The Ghent IAT linked above uses the second approach, with a 200 ms error cross and a 600 ms penalty.
The required-correction version can be built in lab.js by adding a second screen to the trial sequence, after the stimulus screen:
- The screen shows the same stimulus and category labels, plus the red X.
- Its Responses table maps only the correct key, using the same loop column (in the Wisconsin Brief
IAT, the key field is
${ parameters.answer }). Pressing the wrong key again has no effect, so the screen stays until the correct key is pressed. - The screen is skipped when the first response was correct, with a skip condition such as
${ this.state.correct }.
The skip condition needs one additional setting. lab.js normally prepares components before they
run, which means a skip condition that refers to this.state.correct could be evaluated before the
participant has responded to the current stimulus, and would then use the previous trial's result.
Enabling the screen's tardy option delays preparation until just before the screen runs, so the
condition is evaluated after the stimulus response. The Wisconsin Brief IAT sets this option on its
correction screens.
this.state.correct is only updated when a participant gives a response. An IAT without a
response deadline always produces one. If a timeout is later added to the stimulus screen, a trial
that times out leaves the previous trial's value in place, and the correction screen would then
follow that value instead of the current trial.
With this design, each erroneous trial produces two rows in the data: the stimulus screen with the first (incorrect) response, and the correction screen with its own duration. The latency used for scoring is the sum of the two. It is simplest to compute this sum in the analysis script and to keep both rows in the exported data, so that error rates can be reported from the first response.
Counterbalancing block order across participants
Greenwald et al. (2022) describe counterbalancing the order of the two combined pairings as generally desirable, and counterbalancing which side each category's key is on as desirable. Both require assigning each participant to one of several versions of the task.
In Open Lab, this can be done without duplicating the task. Group-code assignment, configured in the
study's landing-page settings, assigns each participant a code (for example, compatible_first and
incompatible_first) with specified probabilities. The code is passed to the lab.js task as the
parameter openlab_group_code, and a script in the task can read it and arrange the blocks in the
corresponding order. The Wisconsin Brief IAT reorders its blocks in a script in the same way, using a
seeded random number instead of an assigned code. To counterbalance key sides as well, four codes
can cover the combinations of order and side.
Two properties of the assignment are relevant here. First, codes are assigned with a minimisation
algorithm (Pocock & Simon, 1975) rather than independent random draws: each new participant is
assigned so that the counts stay close to the specified proportions, which keeps the order groups
close to equal in size even in small samples. Second, participants cannot choose their own group:
link parameters beginning with openlab_ are ignored, so a participant cannot change their
assignment by editing the study URL. The assigned code appears on the Participants page and in the
exported data, where it can be entered as a covariate. In demo mode, a group is assigned at random
on each run, which makes it possible to check that each block order works as intended before
launch.
Timing in the browser
The IAT effect is a difference between reaction times in two sets of blocks, standardised by the participant's own variability. In the lab.js validation study (Henninger et al., 2022), most browsers overestimated reaction times by one to two display frames (roughly 17 to 33 ms), with a standard deviation of up to 7.4 ms. Bridges, Pitiot, MacAskill and Peirce (2020) found that lab.js reaction-time measures had trial-to-trial variability below 9 ms. A lag that is constant within a participant adds the same amount to both pairings and so largely cancels out of the difference. The more relevant practical requirement is the one Knippenberg et al. (2025) stated: a physical keyboard. Greenwald et al. (2022) also recommend reporting whether the task ran in full-screen mode.
For evidence that web-administered IATs are feasible at scale, Project Implicit's public site recorded more than 600,000 completed tasks between October 1998 and April 2000 (Nosek, Banaji, & Greenwald, 2002), and its open Race IAT dataset covers more than 2.3 million participants (Xu, Nosek, & Greenwald, 2014). These are large observational datasets rather than lab-versus-web comparisons. A direct comparison in Carpenter et al. (2019), who ran IATs inside Qualtrics and compared them with the same IAT in Inquisit, found nearly identical results for the two versions.
Scoring the D measure
The D measure (Greenwald, Nosek, & Banaji, 2003) is the standard scoring algorithm. Greenwald et al. (2022, Appendix B) set out the steps for the seven-block IAT:
- Use data from blocks 3, 4, 6 and 7, and discard blocks 1, 2 and 5.
- Remove trials with latencies above 10,000 ms.
- Exclude participants for whom more than 10% of trials are faster than 300 ms.
- If errors required a correction, use the latency to the correct response and add no penalty. Otherwise, replace each error latency with the block mean of correct trials plus 600 ms (or plus two standard deviations).
- For each pair of practice blocks (3 and 6) and test blocks (4 and 7), compute the difference in mean latency and divide it by the standard deviation of all trials in both blocks of that pair. This is an inclusive standard deviation computed across both blocks, not a pooled within-block standard deviation.
- Average the two quotients to obtain D.
For the five-block single-category variant, the same steps apply to the two practice blocks and the two test blocks, as in Knippenberg et al. (2025). For the Brief IAT, Nosek, Bar-Anan, Sriram, Axt and Greenwald (2014) give scoring recommendations. The Wisconsin repository includes an R scoring script for its Brief IAT.
What to report
Greenwald et al. (2022) list what a methods section should contain for an IAT. For an online study, this includes the software and any requirements on the participant's device, whether the task ran full screen, the number of blocks and trials, the interval between trials, the full instructions, how errors were handled, how block order and key sides were counterbalanced, which scoring procedure was used, and an estimate of internal consistency. They also note that internet studies often shorten the task, and that fewer trials reduce reliability. The lab.js task file itself, exported as JSON, can be shared as part of the study materials, as the Ghent group has done.



