• KSEBM
  • Contact us
  • E-Submission
ABOUT
BROWSE ARTICLES
EDITORIAL POLICY
FOR CONTRIBUTORS

Articles

Original Article

Variability, algorithm conformance, and accuracy of large language model–based tools for risk-of-bias (RoB 2.0) assessment of randomized trials: a pilot study

J Evid-Based Pract 2026;2(2):91-97. Published online: September 29, 2026

1Cochrane Korea, Seoul, Korea

2Institute for Evidence-Based Medicine, College of Medicine, Korea University, Seoul, Korea

Correspondence: Hyun Jung Kim Email: iebm.ku@gmail.com
• Received: August 1, 2026   • Accepted: August 31, 2026

© Korean Society of Evidence-Based Medicine, 2026

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 25 Views
  • 0 Download
prev
  • Background
    Large language model (LLM)–based tools are increasingly used to automate risk-of-bias assessment of randomized controlled trials with the revised Cochrane tool (RoB 2.0). However, their reproducibility, their accuracy, and their fidelity to the deterministic algorithm mapping signalling-question responses to domain judgments remain unclear.
  • Methods
    In a conference workshop, participants used an identical prompt and tool versions to assess one RCT with three configurations and entered each tool’s output verbatim (13, 15, and 10 runs). One experienced reviewer’s assessment served as the reference. For each run we computed run-to-run reproducibility, agreement with the expert, and conformance between the tool’s stated domain judgment and the judgment implied by applying the RoB 2.0 algorithm to that run’s own signalling answers.
  • Results
    All configurations showed substantial run-to-run variability under identical conditions, greatest in the conditionally complex domain 2 (Gemini pairwise agreement 0.34). Expert agreement varied widely across configurations (mean 5–60%), and one configuration systematically under-rated risk. Even in domains with a fully specified algorithm, stated judgments frequently diverged from the value implied by the tool’s own signalling answers (domain 2, mean 42%). Skip-logic violations occurred in 69%, 80%, and 100% of runs. Recomputing domain judgments from the signalling answers improved expert agreement for the general-purpose configurations.
  • Conclusion
    LLM-based RoB 2.0 assessments exhibited variability and errors. Whatever tool is adopted, its characteristics and variability must be recognized. Having the tool perform only atomic (single) judgments while delegating aggregation to the algorithm, together with human review, may improve accuracy.
Risk-of-bias assessment of randomized controlled trials (RCTs) is a core step in systematic reviews and in judging the certainty of evidence (GRADE). The revised Cochrane tool, RoB 2.0, has a hierarchical structure: reviewers answer structured signalling questions across five domains, map those answers to a domain-level judgment (low, some concerns, high) via a formal algorithm, and then combine the domain judgments into an overall judgment [1,2]. This process demands substantial time and judgment even from trained assessors and is a bottleneck in large reviews [3,4].
To reduce this burden, automation with large language models (LLMs) has been explored. Following early systems such as RobotReviewer, purpose-built RoB 2.0 tools that combine document retrieval, LLM prompting, and human review—such as ROBoto2—have recently been proposed [5,6]. However, LLMs generate probabilistic output and may respond differently to identical input across runs, and systematic biases arising from alignment training have been reported. How these characteristics affect the reproducibility and accuracy of a task such as RoB 2.0 assessment, which requires multi-step conditional logic and deterministic aggregation, has not been well studied.
Using a design in which a single RCT article was repeatedly assessed with an identical prompt and identical tool versions, we quantified (1) run-to-run reproducibility, (2) accuracy relative to an expert reference, and (3) conformance to the algorithm mapping signalling answers to domain judgments. We further examined whether recomputing domain judgments from a tool’s own signalling answers improves accuracy. We studied three configurations that provide the manual in different ways (full manual in-context, manual via retrieval, and a purpose-built pipeline); rather than ranking these approaches, we focused on characterizing the variability common to all and the error patterns specific to each.
Study Design and Data
At a workshop session of the 3rd conference of the Korean Society for Evidence-Based Medicine, participants used an identical prompt and identical versions of three LLM-based configurations—general-purpose Gemini, a custom Gemini Gem, and ROBoto2 (Roboto)—to assess one RCT article immediately, entering each tool’s output (signalling-question responses and domain judgments) verbatim. Thirteen, fifteen, and ten independent assessments were collected, respectively (Table 1). Because the prompt, tool, and target article were fixed, differences across repeated assessments reflect each configuration’s run-to-run variability.
The three configurations differed in how the RoB 2.0 manual (the detailed tool guidance from the official website) was delivered, and in architecture. Gemini (general) and Gemini Gems used the same base model, the same manual, and the same prompt, differing only in manual delivery: Gemini (general) had the full manual pasted into the chat context so that the entire manual was continuously available to the model, whereas Gemini Gems attached the manual as reference knowledge in a custom Gem. A Gem’s reference knowledge is handled by retrieval-augmented generation (RAG), surfacing only passages relevant to each query. Roboto (ROBoto2) is a purpose-built RoB 2.0 pipeline that combines document retrieval and LLM prompting with the official flowchart aggregation logic. The three configurations thus correspond to full-manual-in-context, manual-via-RAG, and a purpose-built pipeline; our aim was not to rank these approaches but to characterize the variability and error patterns arising in each.
One experienced reviewer’s assessment served as the reference. During initial cross-checking, three transcription errors in the expert’s entries (signalling questions 3.3, 5.1, 5.2) were detected via an algorithm-conformance check and corrected; after correction, the expert’s assessment matched the algorithm-implied judgment in all five domains (Table 1).
RoB 2.0 Algorithm Implementation
We implemented in code the RoB 2.0 algorithm mapping signalling answers to domain judgments. Domains 1, 2, and 5 were fixed using the official guidance decision tables [2]. For domains 3 and 4, the signalling-question structure differs between the 2016 draft and the 2019 final version (the 2016 domain 3 has three questions and domain 4 two, whereas the 2019 final version has four and five, respectively); because our data and the studied tools use the 2019 final structure, we encoded accordingly, and the expert reference matching all five domains provided empirical validation. Typographic variants of the domain-judgment labels were standardized at the analysis stage (raw data unchanged).
Analysis Metrics
Reproducibility. Run-to-run reproducibility per domain was quantified with the Simpson pairwise-agreement index (Σpi2; the probability that two randomly drawn assessments give the same judgment) and normalized Shannon entropy.
Accuracy. The agreement between each tool’s domain judgment and the expert reference was computed per domain.
Algorithm conformance and skip-logic. For each run, we computed the agreement between the judgment implied by applying the algorithm to that run’s own signalling answers and the judgment the tool actually entered. We also counted violations of the skip-logic by which subsequent questions should be marked “not applicable” depending on earlier answers.
Recomputation. Setting aside the tool’s entered judgment, we recomputed the judgment by applying the algorithm to each tool’s own signalling answers and evaluated the change (Δ) in expert agreement.
Statistical Analysis
All analyses were descriptive, consistent with the pilot nature of the study. Domain judgments were coded ordinally (low = 0, some concerns = 1, high = 2). Run-to-run reproducibility was quantified with the Simpson pairwise-agreement index (Σpi2; closer to 1 indicates greater consistency) and normalized Shannon entropy (0 = complete agreement, 1 = maximum dispersion). Because the design is repeated measurement of a single result (one article), Cohen and Fleiss kappa (κ), which presuppose multiple assessment targets, are by definition inappropriate and were not used.
Accuracy relative to the expert and algorithm conformance were computed as domain-level percent agreement. Skip-logic violations were summarized as the number of violations per run and the proportion of runs with ≥1 violation. The recomputation effect was expressed as the change (Δ) obtained by subtracting the expert agreement of the tool’s entered judgment from the expert agreement of the judgment recomputed from the signalling answers.
To convey the precision of the per-configuration summary estimates (five-domain means of reproducibility, expert agreement, and conformance), 95% confidence intervals (CIs) were obtained by nonparametric bootstrap (2,000 resamples, percentile method) resampling runs within each configuration. Formal hypothesis testing of between-configuration differences was not performed, to avoid over-interpretation given the single-target, small-sample pilot design. All analyses were performed in Python (pandas, numpy).
Run-to-Run Reproducibility
Despite identical conditions, all three configurations showed substantial run-to-run variability. The five-domain mean pairwise agreement was 0.76 (95% CI 0.72–0.86) for Roboto, 0.66 (0.61–0.77) for Gems, and 0.59 (0.57–0.69) for Gemini (general) (Table 2; Fig. 1a). Variability differed markedly by domain, being greatest in the conditionally complex domain 2, where Gemini’s pairwise agreement of 0.34 approached chance and Roboto’s was 0.42. Domains 1 and 5, with simpler signalling-question structures, were comparatively stable.
Accuracy Relative to the Expert Reference
Agreement with the expert reference (domain 1 = some concerns; domains 2–5 = high) varied widely across configurations (five-domain means: Roboto 60%, Gemini 24%, Gems 5%; bootstrap 95% CIs 50–70%, 14–35%, 1–11%, respectively; Table 2). In particular, Gems combined moderate reproducibility with only 5% expert agreement, reflecting a tendency to rate most domains as low risk and thus diverging systematically from the expert’s high-risk judgments. Reproducibility (precision) and accuracy (closeness to the reference) were thus independent, and Gems showed a “reproducibly under-rating” directional bias (Fig. 2).
Algorithm Conformance and Internal Consistency
Although the mapping from signalling answers to judgments is deterministic, stated judgments frequently diverged from the implied values even in domains with fully specified official tables. Conformance was 68% for domain 1 and 42% for domain 2 (mean across configurations); Gems was as low as 27% in domain 2 and 13% in domain 4, whereas the expert reference matched in 100% of both domains (Table 3). The five-domain mean conformance was 43% (95% CI 29–58%) for Gemini, 51% (39–63%) for Gems, and 60% (40–78%) for Roboto. This is an error source distinct from stochastic variability, indicating that tools generate plausible signalling answers yet fail to faithfully execute the fixed aggregation rule.
Skip-logic conformance also revealed structural problems. The proportion of runs with ≥1 violation was 69% (mean 2.5 per run) for Gemini, 80% (2.9) for Gems, and 100% (4.3) for Roboto; the expert reference had none (Fig. 1b). Roboto in particular violated the skip-logic in every run, producing answers to all subsequent questions regardless of the preceding results. Moreover, in Roboto some entered judgments did not follow from the tool’s own answers; for example, a run answering “no information” to all three signalling questions of domain 1 had an entered judgment of “low risk,” which no correct algorithm can yield when all answers are “no information” (the official rule gives “some concerns”). Because participants transcribed the tool’s output verbatim, this indicates that the judgment displayed by the tool was not coherent with the answers it displayed—i.e., the answer–judgment link was broken at the tool level.
Effect of Recomputing Judgments from Signalling Answers
Setting aside the tool’s entered judgment, we recomputed the domain judgment by applying the algorithm to each tool’s own signalling answers and assessed the change in expert agreement. For the two general-purpose configurations, whose judgments are generated directly by the LLM, recomputation clearly improved accuracy. Gemini improved in all five domains (mean Δ +0.34; domain 4 +0.54, domain 2 +0.39), and Gems improved substantially in domains 1, 2, and 4 (domain 4 +0.66, domain 2 +0.40; Fig. 3). This suggests that these configurations’ errors arise mainly at the aggregation step rather than in the signalling answers themselves, with the largest gains in the conditionally complex domains 2 and 4.
For Roboto, recomputation tended to lower expert agreement (domains 3, 4, 5), but this result requires caution. Because Roboto’s domain judgment is by design produced by applying the official flowchart to the signalling answers, the very fact that recomputation changes the judgment reflects the answer–judgment decoupling described above. The premise of the recomputation comparison therefore does not hold for this tool, and its recomputation effect is not an interpretable quantity; accordingly, the finding of improved accuracy through recomputation holds only for the configurations that generate judgments directly.
This pilot study, using a repeated-assessment design under identical conditions, identified three limitations of current LLM-based RoB 2.0 assessment.
First, structural limitation. The tools did not faithfully execute RoB 2.0’s conditional logic and deterministic aggregation. Skip-logic violations were pervasive (Roboto violated them in every run), and even in domains with a fully specified official algorithm, entered judgments frequently diverged from the value implied by the tool’s own answers. This reveals a deficit in executing multi-step rules as specified, separate from the ability to generate answers. Notably, Gemini (general) and Gemini Gems were given the same official manual yet showed low conformance and accuracy, indicating that the limitation is not a matter of access to the criteria but of faithfully executing the deterministic multi-step logic.
Second, stochastic variability. Judgments scattered across runs despite identical input, especially in conditionally complex domains (near chance in domain 2). This is a serious constraint for an assessment tool that requires reproducibility.
Third, systematic directional bias. The general-purpose configurations here under-rated risk (Gems expert agreement 5%). The direction of bias, however, depends on the tool and its tuning; a purpose-built pipeline has conversely been reported to be overly conservative [5]. It is therefore more appropriate to say that a systematic directional bias exists and that its direction differs across tools, rather than to generalize a single “optimistic bias.”
Although differing in approach, all three configurations showed substantial run-to-run variability and errors, with somewhat different problem profiles. The configuration providing the full manual in-context (Gemini) had large variability in complex domains and frequently failed at the step of aggregating answers into judgments. The configuration providing the same manual as RAG reference knowledge (Gems) showed a pronounced tendency to under-rate risk, consistent with weaker grounding when only selected passages are surfaced. The purpose-built pipeline (Roboto) had relatively higher reproducibility and accuracy in some domains but showed structural incoherence—violating the skip-logic in every run and displaying judgments that did not follow from its answers. In short, no approach was free of variability and error, and each had its own characteristic weakness.
The implication of this study therefore does not lie in ranking a particular delivery mechanism or architecture. The key point is that, whatever tool is used in the future, its characteristic behavior, variability, and typical error patterns must be adequately recognized and anticipated. When adopting LLM-based RoB assessment, rather than trusting a tool without verification, checking variability through repeated assessment, verifying the consistency between signalling answers and judgments (algorithm conformance), and combining human review with algorithmic aggregation should be standard safeguards.
These findings have practical implications. As the recomputation analysis showed, much of the error in general-purpose configurations arises not in the signalling answers themselves but in the composite step of aggregating them into judgments. Rather than delegating composite, multi-step judgments to the tool, a hybrid approach in which the tool performs only atomic (single) signalling judgments while the deterministic aggregation is handled algorithmically (in code) may improve both accuracy and reproducibility. Indeed, recomputation from signalling answers alone substantially improved domain accuracy for the general-purpose configurations.
Some limitations (the accuracy of the signalling answers) are expected to improve as model capability advances. However, stochastic variability and the faithful execution of deterministic multi-step logic are more robustly addressed by architectural design—separating human-in-the-loop review from algorithmic aggregation—than by awaiting larger models. In addition, our algorithm-conformance check functioned as a data-quality audit, detecting three transcription errors in the expert reference.
Limitations
This is a pilot, hypothesis-generating study based on a single article, a single expert reference, and three configurations (38 assessments in total); generalization requires caution. Given the single-target repeated-measures design, kappa indices were not applied, and the results are specific to particular tool versions and time points. The precise cause of Roboto’s answer–judgment decoupling (implementation, defaults, or display) requires further examination of the tool itself. Larger studies with multiple articles, multiple experts, and various tools and versions are warranted.
Regardless of delivery mechanism or architecture, current LLM-based RoB 2.0 assessment showed stochastic variability, structural non-conformance to deterministic logic, and tool-dependent directional bias, with a distinct error profile for each approach. Although these limitations will be gradually mitigated as AI capability improves, at present the characteristics and variability of whatever tool is used must be heeded. Rather than delegating composite judgments to the tool, it is preferable to focus on the accuracy of single judgments and to combine algorithmic aggregation with human review.

Conflict of Interest

Hyun Jung Kim has been a editor of Journal of Evidence-Based Practice since 2025. However, she was not involved in the peer reviewer selection, evaluation, or decision process of this article. No other potential conflicts of interest relevant to this article were reported.

Funding

The authors received no specific funding for this work.

Data Availability Statement

Not applicable.

Ethics Approval and Consent to Participate

Not applicable.

Authors' Contributions

Conceptualization: HJK. Data curation: HJK JC. Formal analysis: HJK JC HS. Methodology: HJK HS JC. Project administration: MJK HS. Visualization: JC HS MJK. Writing – original draft: HS MJK. Writing – review & editing: HJK JC HS MJK.

Acknowledgments

None.

Fig. 1.
(a) Run-to-run reproducibility (pairwise agreement) by domain. (b) Proportion of runs with ≥1 skip-logic (cascade) violation, by configuration.
jebp-2026-00011f1.jpg
Fig. 2.
Reproducibility (x-axis) versus accuracy relative to the expert reference (y-axis) for each configuration and domain. The upper-right region is ideal; the lower-right region (“reproducibly inaccurate”) is the most concerning.
jebp-2026-00011f2.jpg
Fig. 3.
Change in expert agreement (Δ) when domain judgments are recomputed from each tool’s own signalling answers. Positive values indicate improvement (aggregation error); negative values indicate worsening. For Roboto the negative values reflect answer–judgment decoupling and are of limited interpretability.
jebp-2026-00011f3.jpg
Table 1.
Study Sample and Configurations
Configuration Runs (n) Domain-judgment source Manual delivery
Gemini (general) 13 LLM-generated Full manual in-context
Gemini Gems 15 LLM-generated Manual as RAG knowledge (Gem)
Roboto (ROBoto2) 10 Algorithm (flowchart), as designed Purpose-built RoB 2.0 pipeline
Expert reference 1 Human judgment Algorithm-conformant (5/5, verified)

LLM, large language model; RAG, retrieval-augmented generation. Roboto is designed to produce judgments by applying the official flowchart to the signalling answers.

Table 2.
Domain- Level Reproducibility and Agreement with The Expert Reference, by Configuration
Domain Gemini Gems Roboto
Repro. Agree. Repro. Agree. Repro. Agree.
D1 0.86 8% 0.77 13% 1.00 0%
D2 0.34 38% 0.61 0% 0.42 40%
D3 0.38 38% 0.58 7% 0.54 70%
D4 0.38 38% 0.48 7% 1.00 100%
D5 1.00 0% 0.88 0% 0.82 90%
Mean 0.59 24% 0.66 5% 0.76 60%

Repro., run-to-run reproducibility (Simpson pairwise agreement, 0–1; higher = more consistent). Agree, percent agreement of the domain judgment with the expert reference. D1–D5 = domains 1–5.

Table 3.
Algorithm Conformance, Skip-Logic Violations, and Recomputation Effect, by Configuration
Configuration Conformance (mean) Runs with ≥1 violation Violations/run Recompute Δ (mean)
Gemini (general) 43% 69% 2.5 +0.34
Gemini Gems 51% 80% 2.9 +0.24
Roboto 60% 100% 4.3 Uninterpretable*
Expert reference 100% 0% 0 —

Conformance, proportion of runs in which the entered judgment matched the value implied by the run’s own signalling answers (five-domain mean).

*For Roboto the recomputation effect is not interpretable because of answer–judgment decoupling (see Results 3.4).

  • 1. Higgins JP, Savovic J, Page MJ, Sterne JA. Revised Cochrane risk-of-bias tool for randomized trials (RoB 2) [Internet]. 2019 Aug 22 [cited 2026 08 15]. Available from https://www.riskofbias.info/welcome/rob-2-0-tool/current-version-of-rob-2
  • 2. Sterne JA, Savovic J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 2019; 366: l4898.
  • 3. Crocker TF, Lam N, Jordao M, Brundle C, Prescott M, Forster A, et al. Risk-of-bias assessment using Cochrane's revised tool for randomized trials (RoB 2) was useful but challenging and resource-intensive: observations from a systematic review. J Clin Epidemiol 2023; 161: 39-45.
  • 4. Minozzi S, Cinquini M, Gianola S, Gonzalez-Lorenzo M, Banzi R. The revised Cochrane risk of bias tool for randomized trials (RoB 2) showed low interrater reliability and challenges in its application. J Clin Epidemiol 2020; 126: 37-44.
  • 5. Hevia A, Chintalapati S, Lai VKW, Nguyen TT, Wong WT, Klassen TP, et al. ROBoto2: an interactive system and dataset for LLM-assisted clinical trial risk of bias assessment. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; Stroudsburg (PA): Association for Computational Linguistics; 2025; pp 12-25.
  • 6. Marshall IJ, Kuiper J, Wallace BC. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. J Am Med Inform Assoc 2016; 23: 193-201.

Figure & Data

References

    Citations

    Citations to this article as recorded by  

      Download Citation

      Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

      Format:

      Include:

      Variability, algorithm conformance, and accuracy of large language model–based tools for risk-of-bias (RoB 2.0) assessment of randomized trials: a pilot study
      J Evid-Based Pract. 2026;2(2):91-97.   Published online September 29, 2026
      Download Citation
      Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

      Format:
      • RIS — For EndNote, ProCite, RefWorks, and most other reference management software
      • BibTeX — For JabRef, BibDesk, and other BibTeX-specific software
      Include:
      • Citation for the content below
      Variability, algorithm conformance, and accuracy of large language model–based tools for risk-of-bias (RoB 2.0) assessment of randomized trials: a pilot study
      J Evid-Based Pract. 2026;2(2):91-97.   Published online September 29, 2026
      Close

      Figure

      • 0
      • 1
      • 2
      Variability, algorithm conformance, and accuracy of large language model–based tools for risk-of-bias (RoB 2.0) assessment of randomized trials: a pilot study
      Image Image Image
      Fig. 1. (a) Run-to-run reproducibility (pairwise agreement) by domain. (b) Proportion of runs with ≥1 skip-logic (cascade) violation, by configuration.
      Fig. 2. Reproducibility (x-axis) versus accuracy relative to the expert reference (y-axis) for each configuration and domain. The upper-right region is ideal; the lower-right region (“reproducibly inaccurate”) is the most concerning.
      Fig. 3. Change in expert agreement (Δ) when domain judgments are recomputed from each tool’s own signalling answers. Positive values indicate improvement (aggregation error); negative values indicate worsening. For Roboto the negative values reflect answer–judgment decoupling and are of limited interpretability.
      Variability, algorithm conformance, and accuracy of large language model–based tools for risk-of-bias (RoB 2.0) assessment of randomized trials: a pilot study
      Configuration Runs (n) Domain-judgment source Manual delivery
      Gemini (general) 13 LLM-generated Full manual in-context
      Gemini Gems 15 LLM-generated Manual as RAG knowledge (Gem)
      Roboto (ROBoto2) 10 Algorithm (flowchart), as designed Purpose-built RoB 2.0 pipeline
      Expert reference 1 Human judgment Algorithm-conformant (5/5, verified)
      Domain Gemini Gems Roboto
      Repro. Agree. Repro. Agree. Repro. Agree.
      D1 0.86 8% 0.77 13% 1.00 0%
      D2 0.34 38% 0.61 0% 0.42 40%
      D3 0.38 38% 0.58 7% 0.54 70%
      D4 0.38 38% 0.48 7% 1.00 100%
      D5 1.00 0% 0.88 0% 0.82 90%
      Mean 0.59 24% 0.66 5% 0.76 60%
      Configuration Conformance (mean) Runs with ≥1 violation Violations/run Recompute Δ (mean)
      Gemini (general) 43% 69% 2.5 +0.34
      Gemini Gems 51% 80% 2.9 +0.24
      Roboto 60% 100% 4.3 Uninterpretable*
      Expert reference 100% 0% 0 —
      Table 1. Study Sample and Configurations

      LLM, large language model; RAG, retrieval-augmented generation. Roboto is designed to produce judgments by applying the official flowchart to the signalling answers.

      Table 2. Domain- Level Reproducibility and Agreement with The Expert Reference, by Configuration

      Repro., run-to-run reproducibility (Simpson pairwise agreement, 0–1; higher = more consistent). Agree, percent agreement of the domain judgment with the expert reference. D1–D5 = domains 1–5.

      Table 3. Algorithm Conformance, Skip-Logic Violations, and Recomputation Effect, by Configuration

      Conformance, proportion of runs in which the entered judgment matched the value implied by the run’s own signalling answers (five-domain mean).

      For Roboto the recomputation effect is not interpretable because of answer–judgment decoupling (see Results 3.4).

      TOP