Publication bias is a fundamental threat to the validity of systematic reviews and meta-analyses in clinical medicine. Yet current practice often reduces its assessment to the mechanical application of funnel plots, asymmetry tests, or single adjustment procedures, with limited attention to the underlying assumptions, alternative explanations, or implications for evidence certainty. This narrative methodological article reframes publication bias assessment as an interpretive and editorial responsibility rather than a purely technical problem. We examine what commonly used methods can and cannot reliably support. Detection tools function as nonspecific stress tests that identify deviations from simplified models; they do not diagnose selective publication but highlight situations in which the underlying assumptions require closer inspection. Adjustment approaches, including trim-and-fill, selection models, and regression-based methods, generate hypothetical estimates under unverifiable assumptions**, and therefore provide** sensitivity analyses rather than corrections that recover the true underlying effect. Divergence across adjustment methods is particularly informative, signaling inferential fragility rather than analytical failure. We identify five recurring misinterpretations encountered in peer review: equating asymmetry with proof of publication bias; privileging bias-adjusted estimates as inherently more credible; relying on a single adjustment method without examining assumption dependence; ignoring the plausibility of adjustment direction and magnitude; and overlooking implications for certainty of evidence. Editors and reviewers should prioritize transparency of assumptions, seriously consider alternative explanations, and calibrate conclusions proportionately. Viewing publication bias assessment as an interpretive responsibility rather than a methodological checklist promotes more disciplined inference and strengthens trust in clinical evidence synthesis.
Background Large language model (LLM)–based tools are increasingly used to automate risk-of-bias assessment of randomized controlled trials with the revised Cochrane tool (RoB 2.0). However, their reproducibility, their accuracy, and their fidelity to the deterministic algorithm mapping signalling-question responses to domain judgments remain unclear.
Methods In a conference workshop, participants used an identical prompt and tool versions to assess one RCT with three configurations and entered each tool’s output verbatim (13, 15, and 10 runs). One experienced reviewer’s assessment served as the reference. For each run we computed run-to-run reproducibility, agreement with the expert, and conformance between the tool’s stated domain judgment and the judgment implied by applying the RoB 2.0 algorithm to that run’s own signalling answers.
Results All configurations showed substantial run-to-run variability under identical conditions, greatest in the conditionally complex domain 2 (Gemini pairwise agreement 0.34). Expert agreement varied widely across configurations (mean 5–60%), and one configuration systematically under-rated risk. Even in domains with a fully specified algorithm, stated judgments frequently diverged from the value implied by the tool’s own signalling answers (domain 2, mean 42%). Skip-logic violations occurred in 69%, 80%, and 100% of runs. Recomputing domain judgments from the signalling answers improved expert agreement for the general-purpose configurations.
Conclusion LLM-based RoB 2.0 assessments exhibited variability and errors. Whatever tool is adopted, its characteristics and variability must be recognized. Having the tool perform only atomic (single) judgments while delegating aggregation to the algorithm, together with human review, may improve accuracy.