• KSEBM
  • Contact us
  • E-Submission
ABOUT
BROWSE ARTICLES
EDITORIAL POLICY
FOR CONTRIBUTORS

Articles

Review

Beyond the funnel plot: editorial and reviewer perspectives on publication bias in systematic reviews and meta-analyses in clinical medicine

J Evid-Based Pract 2026;2(2):45-55. Published online: September 29, 2026

Department of Anesthesiology and Pain Medicine, Chung-Ang University College of Medicine, Seoul, Korea

Corresponding authors: Hyun Kang Email: roman00@naver.com
• Received: August 8, 2026   • Revised: August 26, 2026   • Accepted: September 8, 2026

© Korean Society of Evidence-Based Medicine, 2026

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 35 Views
  • 0 Download
prev next
  • Publication bias is a fundamental threat to the validity of systematic reviews and meta-analyses in clinical medicine. Yet current practice often reduces its assessment to the mechanical application of funnel plots, asymmetry tests, or single adjustment procedures, with limited attention to the underlying assumptions, alternative explanations, or implications for evidence certainty. This narrative methodological article reframes publication bias assessment as an interpretive and editorial responsibility rather than a purely technical problem. We examine what commonly used methods can and cannot reliably support. Detection tools function as nonspecific stress tests that identify deviations from simplified models; they do not diagnose selective publication but highlight situations in which the underlying assumptions require closer inspection. Adjustment approaches, including trim-and-fill, selection models, and regression-based methods, generate hypothetical estimates under unverifiable assumptions**, and therefore provide** sensitivity analyses rather than corrections that recover the true underlying effect. Divergence across adjustment methods is particularly informative, signaling inferential fragility rather than analytical failure. We identify five recurring misinterpretations encountered in peer review: equating asymmetry with proof of publication bias; privileging bias-adjusted estimates as inherently more credible; relying on a single adjustment method without examining assumption dependence; ignoring the plausibility of adjustment direction and magnitude; and overlooking implications for certainty of evidence. Editors and reviewers should prioritize transparency of assumptions, seriously consider alternative explanations, and calibrate conclusions proportionately. Viewing publication bias assessment as an interpretive responsibility rather than a methodological checklist promotes more disciplined inference and strengthens trust in clinical evidence synthesis.
Publication bias has long been acknowledged as a fundamental threat to the validity of systematic reviews and meta-analyses in clinical research [1-4]. This refers to the selective processes by which studies with statistically significant or favorable results are more likely to be published, disseminated rapidly, or remain accessible, whereas studies with null or unfavorable findings are delayed, obscured, or never published [1-6].Closely related mechanisms—including time-lag bias, language bias, selective outcome reporting, and dissemination channel bias—operate along the same spectrum and collectively distort the evidentiary landscape available for synthesizing the evidence. The cumulative effect is a systematic inflation of pooled effect estimates and an overstatement of confidence in the resulting conclusions [1-6].
Although publication bias is widely recognized as an important threat to the validity of systematic reviews and meta-analyses, it is often treated as a procedural rather than interpretative issue. During peer review, authors are commonly asked to show that publication bias has been assessed, usually by presenting a funnel plot or reporting the results of a statistical test. These analyses are frequently included as a routine requirement, with little consideration of their assumptions, limitations, alternative explanations for observed findings, or their implications for the certainty of the evidence [7-9]. Although numerous statistical methods have been developed to assess publication bias, ranging from classical funnel plot–based tests to more advanced approaches such as selection models and P curve analysis, there is still no consensus on how their results should be interpreted [1,10-13]. As a result, authors now have access to increasingly sophisticated analytical tools but often lack clear guidance on how to interpret and integrate these findings into balanced conclusions.
Methodological literature on publication bias has largely focused on developing and comparing statistical tools, such as funnel plots, asymmetry tests, and bias-adjustment procedures [1,10-12]. While these methods are indispensable, their application alone does not ensure valid inference. This article contends that the central challenge of publication bias assessment is not methodological sufficiency, but interpretive discipline.
For example, a funnel plot may appear symmetric even when publication bias is present. Conversely, a statistically significant asymmetry test does not necessarily indicate publication bias, as asymmetry may result from factors unrelated to selective publication. Similarly, an adjusted effect estimate should not be considered more credible simply because it is closer to the null value [14,15]. Current evidence suggests that funnel plot asymmetry should be viewed as a nonspecific signal rather than direct evidence of publication bias [14,15]. Asymmetry may arise from genuine differences in treatment effects, variations in study quality, differences in outcome definitions, or differences in follow-up duration, even when no studies are missing. In other words, funnel plot asymmetry indicates that the observed data deviate from an idealized pattern, but it does not identify the reason for that deviation.
This article is organized into four sections. First, we review the available methods for detecting and adjusting for publication bias and clarify what conclusions these methods can support, distinguishing methods that provide sensitivity analyses from those that are often interpreted as diagnostic tests. Second, we discuss five common misinterpretations frequently encountered during peer review and explain why they occur. Third, we propose practical questions that editors and reviewers should consider when evaluating publication bias, with an emphasis on sound interpretation rather than simple methodological compliance. Finally, we discuss publication bias within certainty-of-evidence frameworks, particularly GRADE, and examine situations in which conventional methods provide limited information. Throughout this article, we emphasize that responsible assessment of publication bias requires transparent assumptions, careful consideration of alternative explanations, and conclusions that are consistent with the strength of the available evidence [16-18].
This article does not propose new statistical methods or advocate for any single analytical approach. Rather, it focuses on how existing methods should be interpreted, contextualized, and integrated into judgments about the certainty of evidence. The emphasis is placed on using detection and adjustment methods as interpretive stress tests and integrating their results into transparent, proportionate judgments about evidence certainty—moving beyond checklist compliance toward responsible evidence synthesis.
Publication bias methods are often presented as technical solutions to technical problems. From an editorial perspective, however, their principal value lies not in the numerical estimates they produce but in what they reveal about the limits of inference [19,20]. Therefore, understanding what these methods can—and, just as importantly, cannot —legitimately support is essential for fair, consistent, and proportionate peer review.
Detection tools: stress tests, not diagnostic tools
Detection tools, including funnel plots and small-study effect tests such as Egger’s or Begg’s tests, assess whether observed data are compatible with a simplified reference model—typically one assuming homogeneous effects and the absence of selective dissemination [21,22]. When funnel plot asymmetry or statistically significant small-study effects are observed, the appropriate interpretation is that the reference model may be strained. Crucially, these findings do not identify the mechanism responsible for the deviation.
Consider a meta-analysis of surgical interventions, where smaller trials were conducted at specialized centers with more experienced surgeons, while larger pragmatic trials enrolled diverse centers. The observed asymmetry may reflect genuine effect modification by surgical expertise rather than selective publication. However, reviewers frequently interpret such patterns as confirmatory evidence of missing studies. This example illustrates a fundamental principle: funnel plot asymmetry may arise from true effect modification, methodological gradients, differences in outcome definitions, variations in follow-up duration, or systematic patterns of risk of bias—even when no studies are missing [13,14].
Accordingly, detection results should be interpreted as signals prompting further scrutiny rather than as evidence that publication bias has been demonstrated [21,22]. The appropriate question is not “Has publication bias been proven?” but rather “What might explain the observed pattern, and how much confidence can the evidence reasonably support given this uncertainty?”
Adjustment methods: sensitivity analyses, not corrections
Adjustment methods, including trim-and-fill, selection models, regression-based approaches such as PET-PEESE, and P value–based techniques, extend this logic by imposing explicit assumptions about how selective reporting might operate [23-25]. Each method answers a conditional question: If selection occurred according to a particular mechanism, how would the pooled estimate change? Therefore, the resulting adjusted estimates are hypothetical by construction. They provide information about sensitivity to assumptions but do not recover a true or unbiased underlying effect.
Divergent results across adjustment methods are common because each encodes a different and ultimately unverifiable narrative about the publication process [26].To illustrate, consider how different adjustment methods respond to the same asymmetric funnel plot.
• Trim-and-fill might impute several small negative studies, shifting the pooled estimate toward the null.
• A selection model assuming P value-dependent publication might retain all observed studies but down-weight those with larger standard errors.
• Conversely, PET-PEESE extrapolates to zero precision using regression methods.
Each method encodes a distinct narrative of how selection operates. When these narratives yield divergent estimates, the appropriate inference is not to average them or select the most conservative, but rather to acknowledge that conclusions depend critically on unverifiable assumptions about the publication process. This divergence signals inferential fragility rather than analytical failure.
From an editorial perspective, this divergence is deeply informative. Convergence across methods may suggest relative robustness across multiple assumptions, whereas disagreement signals inferential fragility and warrants restraint in interpretation. What matters is not whether an adjusted estimate is smaller, closer to the null, or statistically non-significant, but rather the transparency with which authors explain the drivers of change and moderate conclusions accordingly [1,27,28].
The distinction between detection, adjustment, and interpretation
In this sense, publication bias methods function most appropriately as interpretive stress tests rather than as corrective algorithms. They expose how sensitive conclusions are to plausible selection mechanisms, modeling choices, and analytical decisions [19,22]. Their primary contribution is not to identify a single preferred estimate but to delineate the range of conclusions that remain defensible under reasonable assumptions. Therefore, editors and reviewers should evaluate publication bias analyses not by their mere presence but by the transparency, coherence, and proportionality with which their implications are interpreted.
Table 1 contrasts these three components. Detection identifies when simplified models are strained, adjustment explores sensitivity to assumptions, and interpretation translates analytical findings into calibrated conclusions. Conflating these stages—for example, treating detection results as proof of publication bias or adjustment outputs as corrections—represents the most consequential category error in publication bias assessment.
Together, these components demonstrate that publication bias assessment is not a linear progression toward a corrected estimate, but rather a process by which analytical findings constrain the range of defensible inferences.
Although publication bias is widely acknowledged as a threat to evidence synthesis, its assessment is frequently misunderstood during the peer review process. These misunderstandings rarely arise from a lack of statistical knowledge. Instead, they reflect a more fundamental problem: the conflation of detection, adjustment, and interpretation as if they were interchangeable steps toward a single, corrected estimate. Predictable errors occur when publication bias methods are treated as diagnostic or corrective tools rather than as components of a broader interpretive process. Five recurring patterns merit particular attention.
Misinterpretation 1: equating asymmetry with proof of publication Bias
Error: One of the most common errors is to equate funnel plot asymmetry or a statistically significant small-study effect test with the presence of publication bias. In peer review, this often appears as statements such as, “Egger’s test is significant; therefore, publication bias exists.” Such reasoning treats detection results as causal diagnoses, overlooking the fact that asymmetry merely indicates a deviation from a simplified reference model assuming homogeneous effects and no selection; it does not explain why that deviation occurred.
Why This Matters: Interpreting asymmetry as proof of selective non-publication collapses multiple plausible explanations into a single unsupported conclusion [1,12]. An evidence base may be fully published and yet still display systematic asymmetry. This misinterpretation leads to an unwarranted downgrading of evidence certainty based on ambiguous signals.
Editorial Response: When asymmetry is detected, editors should ask: “What alternative explanations have the authors considered? Have they examined heterogeneity, methodological quality gradients, and differences in outcome definitions? Are conclusions appropriately cautious given this uncertainty?”
Misinterpretation 2: treating bias-adjusted estimates as inherently more credible
Error: A second critical misinterpretation treats bias-adjusted estimates as inherently more credible than unadjusted results, which is a logical error that misunderstands the hypothetical nature of adjustment procedures. This is often reflected in reviewer requests to present trim-and-fill or similar adjusted results as primary estimates. Such requests disregard the fact that adjustment methods generate estimates under specific unverifiable assumptions about selection.
Why This Matters: An effect estimate that is smaller or closer to the null is not inherently less biased; it may simply reflect the mathematical structure of the chosen method rather than any improvement in evidential validity [1,19,20,23]. Privileging attenuated estimates risks substituting one untested assumption for another rather than improving inference. When an adjusted estimate is smaller, reviewers may feel reassured without scrutinizing whether the implied adjustment is substantively credible.
Editorial Response: Editors should ask authors to present unadjusted and adjusted estimates side by side and to state explicitly which assumption drives the difference between them. An adjusted estimate should be reported as one scenario among several, not as the primary result. Where the adjustment is presented as the headline finding, editors should request that the framing be reversed.
Misinterpretation 3: treating a single adjustment method as definitive
Error: Peer reviewers may focus on whether a publication bias method was applied rather than on how the results behave across multiple plausible approaches. When a single adjusted estimate is presented without a comparison to alternatives, or when divergent results across methods are not discussed, a sensitivity analysis is implicitly elevated to a corrective conclusion. This practice obscures assumption dependence and exaggerates certainty.
Why This Matters: This error is particularly consequential in the presence of substantial between-study heterogeneity, where different adjustment methods encode fundamentally different narratives about selection mechanisms [19,28]. What should function as an exploration of fragility instead becomes a misplaced endpoint, creating false confidence in the single estimate.
Editorial Response: Editors should request that authors present results across multiple adjustment approaches and explicitly discuss how the estimates differ. Stability across diverse methods may support conditional confidence, while divergence signals inferential fragility, which should be reflected in more cautious language and lower certainty-of-evidence ratings.
Misinterpretation 4: ignoring the plausibility of adjustment direction and magnitude
Error: Less frequently articulated but equally consequential is the failure to examine whether the direction and magnitude of an adjustment are substantively plausible. Attenuation toward the null is often accepted as reassuring without scrutiny of whether the implied missing or down-weighted studies resemble realistic trials.
Why This Matters: Consider a meta-analysis of a novel pharmacological intervention in which all included studies were industry-sponsored and showed large effects. Trim-and-fill analysis may impute several small negative studies to achieve symmetry. However, if no such negative studies exist in the evidence base despite extensive searching, the imputed studies may be implausibly pessimistic. Adjustments that shift influence toward studies unlike any observed in the literature should prompt concern rather than confidence [19,23].
Editorial Response: When adjustment methods are employed, editors should verify that (1) the implied selection mechanism reflects plausible publication dynamics in the field, (2) hypothetical imputed studies resemble realistic trials, and (3) the adjustment does not systematically disadvantage methodologically rigorous but smaller studies.
Misinterpretation 5: overlooking implications for certainty of evidence
Error: Finally, publication bias analyses are often discussed in isolation without explicit consequences for interpretation. Even when adjusted estimates vary widely or demonstrate strong dependence on modeling assumptions, manuscripts may retain confident causal language and strong clinical recommendations. In such cases, the central issue is not whether publication bias was “detected,” but whether its potential impact has been appropriately translated into judgments about the certainty of evidence and the strength of conclusions [20,28-30].
Why This Matters: This disconnection between analytical findings and interpretive restraint represents a fundamental breakdown in the evidence synthesis discipline. When the analysis reveals inferential fragility, the conclusions must correspondingly weaken. Maintaining strong recommendations despite demonstrated sensitivity to unverifiable assumptions violates the principle of proportionality [29,32].
Editorial Response: Editors should assess whether strong assumption dependence, wide sensitivity ranges, or instability across models are reflected in more cautious language, explicit downgrading of certainty of evidence per GRADE framework, and moderated recommendations. When analytical fragility is documented but conclusions remain confident, an unresolved disconnect persists between the analysis and interpretation.
Summary: These five recurring misinterpretations highlight the need for an explicit editorial framework that distinguishes detection from adjustment and both from interpretation. Shifting attention away from method selection and toward inferential judgment, transparency, and proportional interpretation is essential if publication bias analyses are to meaningfully inform evidence appraisal rather than merely satisfying procedural expectations.
From an editorial perspective, the central task in evaluating publication bias analyses is not to verify whether specific statistical procedures have been applied but to determine whether the analysis demonstrates sound inferential reasoning. This requires shifting attention away from method labels and toward the logic that connects the assumptions, analytical results, and resulting conclusions. In practice, these questions function as safeguards against the category errors outlined above, particularly the conflation of detection, adjustment, and interpretation [19,31].
What assumptions does the analysis rely on, and are they explicitly stated?
All publication bias methods encode assumptions regarding how studies are selected, reported, or disseminated. Therefore, editorial evaluation should begin by asking whether these assumptions are clearly articulated and whether they plausibly reflect the clinical and publication contexts of the evidence base. Analyses that present bias-adjusted estimates without explaining the implied selection mechanism provide little interpretive grounding regardless of statistical sophistication. Without explicit assumptions, neither the detection nor the adjustment results can be meaningfully interpreted, and any downstream conclusions remain unsupported [4,20,31].
Are alternative explanations for the observed patterns seriously considered?
When funnel plot asymmetry or small-study effects are reported, authors should address explanations beyond selective non-publication. These include true effect modification, gradients in methodological quality, differences in outcome definitions, variations in follow-up duration, and analytical flexibility. Editorial judgment should favor analyses that situate publication bias within this broader explanatory landscape rather than treating it as a default diagnosis or a standalone issue. Failure to consider alternatives transforms detection signals into causal claims, reinforcing the misinterpretations commonly encountered in peer reviews [1,11,19,32].
How stable are the conclusions across reasonable analytical choices?
A robust publication bias assessment makes sensitivity visible. Editors should look for transparent comparisons across multiple plausible adjustment methods or model specifications, including scenarios in which the conclusions weaken, attenuate, or change direction. Stability across approaches may support conditional confidence, whereas divergence indicates inferential fragility. Crucially, both outcomes are informative only if they are reported and interpreted and not suppressed in favor of a single adjusted estimate [20,28,32].
Have the direction and magnitude of the adjustment been critically examined?
Noting that an adjusted estimate is smaller or closer to the null hypothesis is insufficient. Editorial scrutiny should consider whether implied adjustments are substantively credible. Do the hypothetical missing or down-weighted studies resemble plausible trials in the field? Does the adjustment disproportionately privilege larger but lower-quality studies over smaller, methodologically rigorous ones? Unexamined attenuation toward the null risks substituting one untested assumption for another rather than improving inference [20,28,32].
Are the consequences for the certainty of evidence clearly articulated?
Most importantly, publication bias analysis should have visible consequences for interpretation. Editors should assess whether strong assumption dependence, wide sensitivity ranges, or instability across models are reflected in more cautious language, explicit downgrading of the certainty of evidence, or moderated recommendations. When analytical fragility is documented but conclusions remain confident, an unresolved disconnect persists between analysis and interpretation. At this stage, publication bias assessment ceases to be a methodological issue and becomes a question of evidential responsibility [30,33].
These questions emphasize that publication bias assessment is fundamentally an exercise in interpretation rather than algorithm selection. Editorial evaluation should therefore prioritize transparency, restraint, and coherence of reasoning over mechanical compliance with the methodological checklists. Viewed through this lens, publication bias analysis directly informs judgments about the certainty of evidence rather than functioning as a stand-alone diagnostic step or a means of effect size correction.
Within evidence-based medicine, publication bias is best understood not as a statistical inconvenience but as a fundamental threat to the certainty of evidence. Frameworks such as GRADE explicitly recognize publication bias as a reason to downgrade confidence in effect estimates, reflecting the insight that selective dissemination distorts the evidentiary base itself rather than merely shifting numerical results [18,33]. From this perspective, the goal of publication bias assessment is not to repair an estimate but to determine how cautiously that estimate should be interpreted.
Publication bias as evidence fragility
This reframing has direct implications for editorial decision making. Inferential fragility is indicated when bias-adjusted estimates vary widely across plausible assumptions, results are highly sensitive to unverifiable selection mechanisms, or conclusions hinge on a small number of influential studies. In GRADE terms, such fragility justifies downgrading the certainty of evidence, even when statistical significance persists under the selected models. The apparent numerical robustness within a single analytical framework does not override the instability across reasonable alternatives [18,33].
Equally important, the absence of statistically significant evidence of publication bias should not be interpreted as reassurance. Funnel plot symmetry, non-significant asymmetry tests, or convergence of a single adjustment method merely indicate that selective processes were not detected under the assumptions applied. They do not establish that publication bias is absent. Treating “no evidence of bias” as evidence of no bias conflates detection with interpretation, which is a mistake particularly consequential in settings characterized by substantial heterogeneity, flexible outcome definitions, or complex reporting practices, where conventional detection tools are known to be weakly informative [7,12,18].
When evidence can retain higher certainty
Conversely, when publication bias analyses demonstrate relative stability across multiple methods and assumptions, especially in evidence bases with prespecified outcomes, low heterogeneity, and external constraints such as trial registration, editors may reasonably judge the risk of publication bias to be lower. However, even under these circumstances, the appropriate inference is conditional confidence rather than confirmation of a true effect. Certainty may be retained, but it remains explicitly contingent on the assumptions examined[19,21].
Integration with GRADE framework
Viewed through a certainty-of-evidence lens, publication bias analysis functions as the final interpretive step linking statistical modeling to editorial responsibility. Its primary contribution lies in translating analytical sensitivity into calibrated language, moderated recommendations, and explicit acknowledgment of uncertainties. At this stage, publication bias assessment ceases to be a methodological exercise and becomes an ethical one, requiring that claims be proportionate to the robustness of the evidence supporting them [18,33].
In this sense, publication bias analysis does not tell us which estimate to trust; it tells us how much trust the evidence can reasonably bear.
Although publication bias is a general concern in all forms of evidence synthesis, certain research contexts pose a heightened risk of misinterpretation. In these settings, conventional detection or adjustment methods may appear reassuring while obscuring deeper forms of selective dissemination and analytical flexibility. For editors and reviewers, the challenge is not the absence of methods but the presence of apparently coherent results in contexts where inferential fragility is structurally amplified [15,34] (Table 2).
Umbrella reviews and higher-order syntheses
Umbrella reviews synthesize evidence from multiple meta-analyses and therefore operate at an increased distance from the primary study selection. This layered structure magnifies the consequences of selective dissemination, rather than neutralizing them. Publication and reporting biases at the primary study level may propagate through individual meta-analyses and be further reinforced when umbrella reviews preferentially cite statistically significant or highly cited syntheses [35,36].
In addition, the overlap of primary studies across meta-analyses creates an illusion of independent replication, as the same influential trials—often those with large or favorable effects—are repeatedly reintroduced. In such settings, funnel plot–based assessments conducted at the meta-analysis level may appear well-behaved even when the underlying evidence base is highly selective, requiring editorial scrutiny that extends beyond per-meta-analysis diagnostics [35,36].
Flexible outcomes, time points, and analytical choices
Publication bias is commonly conceptualized as the non-publication of entire studies. However, in many clinical studies, the dominant selective process operates within studies rather than between them. Selective reporting of outcomes, time points, subgroups, or analytical models can generate robust pooled effects, even when all studies are published [13,36].
Conventional publication bias methods are poorly equipped to detect such within-study selection, and funnel plot symmetry or stable-adjusted estimates may therefore provide false reassurance. Editors should be particularly alert when outcome definitions are heterogeneous, follow-up windows are flexible, or analytical decisions are weakly prespecified, as these features expand the space for selective reporting [15,37,38].
Preprints, living reviews, and rapidly evolving literatures
The increasing use of preprints and living systematic reviews has altered the dynamics of the selective dissemination of research. The inclusion of preprints may mitigate traditional publication bias by capturing studies that have not yet been filtered by journal publication. Simultaneously, rapid dissemination environments may preferentially amplify novel, positive, or attention-grabbing findings, introducing new forms of selection.
Editorial evaluation should therefore focus not on whether preprints are included but on how their inclusion affects interpretability. Analyses that explicitly examine the influence of preprints on pooled estimates provide more informative evidence than their unexamined inclusion or exclusion [31,39].
Automation and AI-assisted evidence synthesis
Automation and artificial intelligence–assisted tools are increasingly used for literature screening, prioritization, and synthesis. While these tools improve efficiency, they may also inherit and amplify existing publication and citation biases. Algorithms trained predominantly on published or highly cited studies risk reinforcing selective visibility, rather than correcting it.
Transparency is essential from an editorial perspective. Editors and reviewers should assess whether the automated processes were audited, whether the screening and prioritization decisions were reproducible, and whether human judgment was retained at the interpretive stage, particularly in the publication bias assessment [40,41].
Small evidence bases and influential trials
Finally, publication bias methods are especially fragile when the number of included studies is small or when a few trials dominate the evidence. In such circumstances, both detection and adjustment procedures can behave erratically, with single influential studies driving apparent asymmetry or large shifts in the adjusted estimates. Analytical sophistication cannot compensate for the lack of evidence.
Therefore, editorial judgment should favor the explicit acknowledgment of uncertainty over reliance on any single bias-adjustment procedure [1,15].Transparency about these limitations is more informative than false confidence in methods applied to inadequate evidence bases.
These contexts underscore that publication bias assessment cannot be interpreted independently of the evidence base’s structure and dynamics. In complex, flexible, or rapidly evolving literatures, restraint, transparency, and attention to the certainty of evidence are often more informative than elaborate modeling. For editors and reviewers, recognizing when methods are weakly informative is as important as their correctness.
Rather than prescribing specific statistical methods, the following checklist emphasizes interpretive quality, transparency and inferential discipline. Each domain reflects a question that editors and reviewers should consider when assessing whether publication bias analyses meaningfully inform the interpretation.
This checklist is intentionally brief and interpretive in nature. Its purpose is not to standardize which methods must be used but to ensure that how methods are used aligns with what they can legitimately support.
Publication bias methods do not correct bias in a causal sense. When used well, they reveal how dependent clinical conclusions are on unverifiable assumptions about selection, reporting, and study visibility. Divergent adjusted estimates, wide sensitivity ranges, and strong dependence on modeling choices are not analytical failures; they are signals of evidential fragility. Recognizing this distinction transforms publication bias assessment from a technical problem into an interpretive responsibility.
For authors, responsible use of publication bias methods means resisting the temptation to present a single “corrected” estimate and instead documenting how conclusions change across plausible scenarios. For reviewers, this means evaluating coherence and transparency rather than technical box-checking. For editors, this means recognizing that the most important question is not whether publication bias was tested but whether uncertainty was honestly carried through to interpretation and claims.
Viewed through a certainty-of-evidence framework such as GRADE, the primary contribution of publication bias analysis is to inform the appropriate downgrading of confidence. When conclusions hinge on strong, unverifiable assumptions, certainty should be reduced accordingly, even if numerical estimates appear precise. In this sense, publication bias analysis serves as a bridge between statistical modeling and responsible interpretation of the results.
Treating publication bias as an interpretive responsibility rather than a statistical obstacle helps ensure that meta-analyses support clinical recommendations that are proportional to the strength of the underlying evidence. This shift from correction to calibration is essential for preserving trust in evidence synthesis in an era of increasing analytical complexity and automation. By prioritizing transparency, coherence, and proportional inference over mechanical compliance with analytical checklists, editors and reviewers can ensure that publication bias assessment meaningfully informs evidence appraisal and strengthens, rather than undermines, confidence in clinical evidence synthesis.

Conflict of Interest

Hyun Kang has been the Editor-in-Chief of the Journal of Evidence-Based Practice since 2025; however, he was not involved in the peer reviewer selection, evaluation, or decision process of this article. No other potential conflicts of interest relevant to this article were reported.

Funding

This research was funded by the Basic Science Research Program through the National Research Foundation (NRF) of Korea, funded by the Ministry of Education, Science, and Technology (NRF-2022R1F1A1074934).

Data Availability Statement

Data sharing is not applicable to this article as no datasets were generated or analyzed during the current study.

Ethics Approval and Consent to Participate

Not applicable.

Authors' Contributions

Conceptualization: HK. Funding acquisition: HK. Methodology: HK. Writing – original draft: HK. Writing – review & editing: HK.

Acknowledgments

None.

Table 1.
Detection, Adjustment, and Interpretation in the Editorial Evaluation of Publication Bias
Dimension Detection Methods Adjustment Methods Interpretation (Editorial Task)
Primary purpose Identify deviations from a simplified reference model Explore how pooled estimates change under assumed selection mechanisms Judge how much confidence the evidence can reasonably support
Typical examples Funnel plots; Egger’s test; Begg’s test Trim-and-fill; selection models; PET-PEESE; P value–based methods Editorial synthesis of bias analyses, heterogeneity, and study quality
Core question addressed Do the observed data fit a model with no selection or homogeneous effects? If selection operates in a specific way, how would the estimate change? How should conclusions be calibrated, given assumption dependence and inferential fragility?
Nature of output Graphical patterns or test statistics Hypothetical bias-adjusted effect estimates Judgments about certainty of evidence and strength of claims
Key assumptions Homogeneous effects; absence of selective dissemination Explicit, method-specific selection mechanisms Transparency, plausibility, and proportionality of inference
What the results can tell us That the simple reference model may be violated That conclusions are sensitive (or robust) to particular assumptions Whether confidence should be downgraded or conclusions moderated
What the results cannot tell us The cause of asymmetry or deviation The true underlying effect size A single ‘corrected’ or definitive estimate
Common misinterpretation Asymmetry interpreted as proof of publication bias Adjusted estimate treated as closer to the truth Confidence maintained despite demonstrated fragility
Appropriate editorial use Prompt further scrutiny and exploration of alternatives Function as sensitivity analyses, not replacements Integrate findings into certainty-of-evidence judgments (e.g., GRADE)
Risk if overemphasized Overdiagnosis of publication bias False sense of correction or precision Overconfident conclusions despite uncertainty

PET-PEESE: precision-effect test and precision-effect estimate with standard error, GRADE: Grading of Recommendations, Assessment, Development, and Evaluations.

Table 2.
Domain and Key Questions Editors and Reviewers Should Ask
Domain Key Questions Editors and Reviewers Should Ask
Conceptual framing Do the authors clearly explain what publication bias methods can and cannot tell us? Are adjustment methods framed as sensitivity analyses rather than as causal corrections?
Detection Are funnel plots or small-study effect tests interpreted as stress tests rather than proof of publication bias? Are alternative explanations, such as heterogeneity or risk-of-bias gradients, discussed?
Choice of methods Is the choice of publication bias methods justified in relation to plausible selection mechanisms in the field, rather than being applied mechanically?
Multiplicity of models Were multiple adjustment approaches presented when publication bias was suspected? If the results differ, is this divergence highlighted and interpreted?
Transparency of assumptions Do the authors clearly describe the assumptions encoded by each method (e.g., symmetry, P value dependence, precision–effect relationships)?
Integration with heterogeneity Are publication bias analyses interpreted alongside heterogeneity, risk of bias assessments, and information size, rather than in isolation?
Certainty of evidence Are the findings from publication bias analyses explicitly linked to judgments about the certainty of the evidence and the strength of the conclusions?
Scope and limits Do the authors acknowledge when publication bias methods are weakly informative, such as when few studies or substantial heterogeneity are present
Reporting balance Are both unadjusted and adjusted results presented transparently, without privileging a single ‘corrected’ estimate?
Conclusion discipline Do the conclusions appropriately reflect evidential fragility when the results are sensitive to assumptions, or do they retain strong causal language despite uncertainty?
  • 1. Debray TPA, Moons KGM, Riley RD. Detecting small-study effects and funnel plot asymmetry in meta-analysis of survival data: A comparison of new and existing tests. Res Synth Methods 2018; 9: 41-50.
  • 2. Sterne JA, Egger M, Smith GD. Systematic reviews in health care: investigating and dealing with publication and other biases in meta-analysis. BMJ 2001; 323: 101-5.
  • 3. Dwan K, Altman DG, Arnaiz JA, Bloom J, Chan AW, Cronin E, et al. Systematic review of the empirical evidence of study publication bias and outcome reporting bias. PLoS One 2008; 3: e3081.
  • 4. Hopewell S, Loudon K, Clarke MJ, Oxman AD, Dickersin K. Publication bias in clinical trials due to statistical significance or direction of trial results. Cochrane Database Syst Rev 2009; 1: MR000006.
  • 5. Marks-Anglin A, Chen Y. A historical review of publication bias. Res Synth Methods 2020; 11: 725-42.
  • 6. Cheema HA, Shahid A, Ehsan M, Ayyan M. The misuse of funnel plots in meta-analyses of proportions: are they really useful? Clin Kidney J 2022; 15: 1209-10.
  • 7. Afonso J, Ramirez-Campillo R, Clemente FM, Buttner FC, Andrade R. The perils of misinterpreting and misusing "publication bias" in meta-analyses: an education review on funnel plot-based methods. Sports Med 2024; 54: 257-69.
  • 8. Hamilton DG, Fraser H, Hoekstra R, Fidler F. Journal policies and editor's opinions on peer review. eLife 2020; 9: e62529.
  • 9. Page MJ, Higgins JP, Sterne JA. Assessing risk of bias due to missing results in a synthesis. editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. In: Higgins JP, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, , London, Cochrane. 2024.
  • 10. Sterne JA, Egger M. Funnel plots for detecting bias in meta-analysis: guidelines on choice of axis. J Clin Epidemiol 2001; 54: 1046-55.
  • 11. Zwetsloot PP, Van Der Naald M, Sena ES, Howells DW, IntHout J, De Groot JA, et al. Standardized mean differences cause funnel plot distortion in publication bias assessments. Elife 2017; 6: e24260.
  • 12. Doleman B, Freeman SC, Lund JN, Williams JP, Sutton AJ. Funnel plots may show asymmetry in the absence of publication bias with continuous outcomes dependent on baseline risk: presentation of a new publication bias test. Res Synth Methods 2020; 11: 522-34.
  • 13. Kang H. Correcting what cannot be corrected: rethinking publication bias analysis methods in clinical meta-analyses. Korean J Anesthesiol 2026; 79: 271-90.
  • 14. Muradchanian J, Hoekstra R, Kiers H, van Ravenzwaaij D. The role of results in deciding to publish: A direct comparison across authors, reviewers, and editors based on an online survey. PLoS One 2023; 18: e0292279.
  • 15. Mathur MB, VanderWeele TJ. Estimating publication bias in meta-analyses of peer-reviewed studies: A meta-meta-analysis across disciplines and journal tiers. Res Synth Methods 2021; 12: 176-91.
  • 16. Schünemann HJ, Neumann I, Hultcrantz M, Brignardello-Petersen R, Zeng L, Murad MH, et al. GRADE guidance 35: update on rating imprecision for assessing contextualized certainty of evidence and making decisions. J Clin Epidemiol 2022; 150: 225-42.
  • 17. Higgins JP, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. London, Cochrane. 2024.
  • 18. Prasad M. Introduction to the GRADE tool for rating certainty in evidence and recommendations. Clin Epidemiol Glob Health 2024; 25: 101484.
  • 19. Mavridis D, Salanti G. How to assess publication bias: funnel plot, trim-and-fill method and selection models. Evid Based Ment Health 2014; 17: 30.
  • 20. Lin L, Chu H. Quantifying publication bias in meta-analysis. Biometrics 2018; 74: 785-94.
  • 21. Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ 1997; 315: 629-34.
  • 22. Møller AP, Jennions MD. Testing and adjusting for publication bias. Trends Ecol Evol 2001; 16: 580-6.
  • 23. Shi L, Lin L. The trim-and-fill method for publication bias: practical guidelines and recommendations based on a large database of meta-analyses. Medicine (Baltimore) 2019; 98: e15987.
  • 24. Maier M, VanderWeele TJ, Mathur MB. Using selection models to assess sensitivity to publication bias: A tutorial and call for more routine use. Campbell Syst Rev 2022; 18: e1256.
  • 25. Alinaghi N, Reed WR. Meta-analysis and publication bias: how well does the FAT-PET-PEESE procedure work? Res Synth Methods 2018; 9: 285-311.
  • 26. Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Comparison of two methods to detect publication bias in meta-analysis. JAMA 2006; 295: 676-80.
  • 27. Bartoš F, Maier M, Quintana DS, Wagenmakers E-J. Adjusting for publication bias in JASP and R: selection models, PET-PEESE, and robust Bayesian meta-analysis. Adv Methods Pract Psychol Sci 2022; 5: 25152459221109259.
  • 28. Ahmed I, Sutton AJ, Riley RD. Assessment of publication bias, selection bias, and unavailable data in meta-analyses using individual participant data: a database survey. BMJ 2012; 344: d7762.
  • 29. Mathur MB. Assessing robustness to worst case publication bias using a simple subset meta-analysis. BMJ 2024; 384: e076851.
  • 30. Al Duhailib Z, Granholm A, Alhazzani W, Oczkowski S, Belley-Cote E, Møller MH. GRADE pearls and pitfalls-part 1: systematic reviews and meta-analyses. Acta Anaesthesiol Scand 2024; 68: 584-92.
  • 31. Page MJ, Sterne JAC, Higgins JPT, Egger M. Investigating and dealing with publication bias and other reporting biases in meta-analyses of health research: A review. Res Synth Methods 2021; 12: 248-59.
  • 32. Formann AK. Estimating the proportion of studies missing for meta-analysis due to publication bias. Contemp Clin Trials 2008; 29: 732-9.
  • 33. Daou JP, Riera R, Pacheco RL. Reasons for downgrading the certainty of evidence for indirectness in synthesis of surgical procedures for patients with fractures: a meta-research analysis. J Eval Clin Pract 2025; 31: e70091.
  • 34. Ropovik I, Adamkovic M, Greger D. Neglect of publication bias compromises meta-analyses of educational research. PLoS One 2021; 16: e0252415.
  • 35. Fusar-Poli P, Radua J. Ten simple rules for conducting umbrella reviews. Evid Based Ment Health 2018; 21: 95-100.
  • 36. Choi GJ, Kang H. The umbrella review: a useful strategy in the rain of evidence. Korean J Pain 2022; 35: 127-8.
  • 37. Song F, Parekh S, Hooper L, Loke YK, Ryder J, Sutton AJ, et al. Dissemination and publication of research findings: an updated review of related biases. Health Technol Assess 2010; 14: iii, ix-xi, 1-193.
  • 38. Lindsley K, Fusco N, Li T, Scholten R, Hooft L. Clinical trial registration was associated with lower risk of bias compared with non-registered trials among trials included in systematic reviews. J Clin Epidemiol 2022; 145: 164-73.
  • 39. Kang H. Beyond the paywall: the role of preprints in overcoming publication bias. J Evid-Based Pract 2025; 1: 7-11.
  • 40. Moens M, Nagels G, Wake N, Goudman L. Artificial intelligence as team member versus manual screening to conduct systematic reviews in medical sciences. iScience 2025; 28: 113559.
  • 41. Millard LA, Flach PA, Higgins JP. Machine learning to assist risk-of-bias assessments in systematic reviews. Int J Epidemiol 2016; 45: 266-77.

Figure & Data

References

    Citations

    Citations to this article as recorded by  

      Download Citation

      Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

      Format:

      Include:

      Beyond the funnel plot: editorial and reviewer perspectives on publication bias in systematic reviews and meta-analyses in clinical medicine
      J Evid-Based Pract. 2026;2(2):45-55.   Published online September 29, 2026
      Download Citation
      Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

      Format:
      • RIS — For EndNote, ProCite, RefWorks, and most other reference management software
      • BibTeX — For JabRef, BibDesk, and other BibTeX-specific software
      Include:
      • Citation for the content below
      Beyond the funnel plot: editorial and reviewer perspectives on publication bias in systematic reviews and meta-analyses in clinical medicine
      J Evid-Based Pract. 2026;2(2):45-55.   Published online September 29, 2026
      Close
      Beyond the funnel plot: editorial and reviewer perspectives on publication bias in systematic reviews and meta-analyses in clinical medicine
      Beyond the funnel plot: editorial and reviewer perspectives on publication bias in systematic reviews and meta-analyses in clinical medicine
      Dimension Detection Methods Adjustment Methods Interpretation (Editorial Task)
      Primary purpose Identify deviations from a simplified reference model Explore how pooled estimates change under assumed selection mechanisms Judge how much confidence the evidence can reasonably support
      Typical examples Funnel plots; Egger’s test; Begg’s test Trim-and-fill; selection models; PET-PEESE; P value–based methods Editorial synthesis of bias analyses, heterogeneity, and study quality
      Core question addressed Do the observed data fit a model with no selection or homogeneous effects? If selection operates in a specific way, how would the estimate change? How should conclusions be calibrated, given assumption dependence and inferential fragility?
      Nature of output Graphical patterns or test statistics Hypothetical bias-adjusted effect estimates Judgments about certainty of evidence and strength of claims
      Key assumptions Homogeneous effects; absence of selective dissemination Explicit, method-specific selection mechanisms Transparency, plausibility, and proportionality of inference
      What the results can tell us That the simple reference model may be violated That conclusions are sensitive (or robust) to particular assumptions Whether confidence should be downgraded or conclusions moderated
      What the results cannot tell us The cause of asymmetry or deviation The true underlying effect size A single ‘corrected’ or definitive estimate
      Common misinterpretation Asymmetry interpreted as proof of publication bias Adjusted estimate treated as closer to the truth Confidence maintained despite demonstrated fragility
      Appropriate editorial use Prompt further scrutiny and exploration of alternatives Function as sensitivity analyses, not replacements Integrate findings into certainty-of-evidence judgments (e.g., GRADE)
      Risk if overemphasized Overdiagnosis of publication bias False sense of correction or precision Overconfident conclusions despite uncertainty
      Domain Key Questions Editors and Reviewers Should Ask
      Conceptual framing Do the authors clearly explain what publication bias methods can and cannot tell us? Are adjustment methods framed as sensitivity analyses rather than as causal corrections?
      Detection Are funnel plots or small-study effect tests interpreted as stress tests rather than proof of publication bias? Are alternative explanations, such as heterogeneity or risk-of-bias gradients, discussed?
      Choice of methods Is the choice of publication bias methods justified in relation to plausible selection mechanisms in the field, rather than being applied mechanically?
      Multiplicity of models Were multiple adjustment approaches presented when publication bias was suspected? If the results differ, is this divergence highlighted and interpreted?
      Transparency of assumptions Do the authors clearly describe the assumptions encoded by each method (e.g., symmetry, P value dependence, precision–effect relationships)?
      Integration with heterogeneity Are publication bias analyses interpreted alongside heterogeneity, risk of bias assessments, and information size, rather than in isolation?
      Certainty of evidence Are the findings from publication bias analyses explicitly linked to judgments about the certainty of the evidence and the strength of the conclusions?
      Scope and limits Do the authors acknowledge when publication bias methods are weakly informative, such as when few studies or substantial heterogeneity are present
      Reporting balance Are both unadjusted and adjusted results presented transparently, without privileging a single ‘corrected’ estimate?
      Conclusion discipline Do the conclusions appropriately reflect evidential fragility when the results are sensitive to assumptions, or do they retain strong causal language despite uncertainty?
      Table 1. Detection, Adjustment, and Interpretation in the Editorial Evaluation of Publication Bias

      PET-PEESE: precision-effect test and precision-effect estimate with standard error, GRADE: Grading of Recommendations, Assessment, Development, and Evaluations.

      Table 2. Domain and Key Questions Editors and Reviewers Should Ask

      TOP