📖 All StatFacts guides

Multi-Armed Bandits vs. Fixed-Horizon A/B Tests: Choosing an Allocation Strategy

Bandits and fixed-horizon A/B tests optimize for different things — cumulative reward during the test versus a clean, comparable effect estimate afterward. Here's how to decide which one fits your decision, and how to read StatFacts benchmarks against results from either.

Bayesian vs. Frequentist A/B Testing: Choosing the Right Framework

Frequentist p-values and Bayesian posterior probabilities answer different questions about the same experiment. Here's how to pick the right one and read StatFacts benchmarks correctly under each.

Stopping Rules for A/B Tests: How to Avoid the Peeking Problem

Checking a test's p-value every day and stopping the moment it looks significant quietly inflates your false positive rate far past 5%. This guide walks through fixed-horizon, sequential, and alpha-spending stopping rules so you can call winners without fooling yourself.

Heterogeneous Treatment Effects: How to Read Subgroup Lift Without Fooling Yourself

Average treatment effects hide as much as they reveal — this guide covers how to analyze segment-level lift responsibly, from pre-registration to shrinkage to knowing when a subgroup difference is real. It's written for teams using effect-size benchmarks to decide whether a personalization strategy is worth building.

How to Calculate Minimum Detectable Effect (MDE) for Effective Power Analysis

Understanding Minimum Detectable Effect (MDE) is crucial for designing statistically powerful experiments, ensuring you allocate appropriate resources to detect meaningful changes. This guide provides a practical methodology for calculating MDE, enabling product and growth teams to plan efficient A/B tests and data-driven initiatives.

Safeguarding Insights: Preventing P-Hacking and Data Dredging in Post-Hoc Analysis

Understand the insidious risks of p-hacking and data dredging, which inflate false positives and erode trust in analytics. This guide provides practical, ethical methodologies for rigorous post-hoc analysis, leveraging established benchmarks to deliver reliable insights for your team.

Validating Your Experimentation Platform with A/A Testing: A Practical Guide

A/A testing is a critical step to ensure your experimentation platform is trustworthy before running A/B tests. This guide outlines a methodical approach to validate data integrity and metric stability.

Choosing the correct unit of randomization in experiments

Choosing the correct unit of randomization in experiments

Optimizing A/B Tests for Low-Traffic Sites: A Practical Methodology for Startups

For startups and small teams with limited traffic, traditional A/B testing methods often struggle to yield significant results quickly. This guide provides a practical methodology to run effective A/B tests on low-traffic sites and products by leveraging benchmarks, heuristics, and strategic prioritization to optimize conversion rates.

Navigating One-tailed vs. Two-tailed Hypothesis Tests in Product Experiments

Selecting between one-tailed and two-tailed hypothesis tests is crucial for product managers and analysts, directly impacting an experiment's statistical power and the validity of conclusions drawn from A/B test results.

Boosting A/B Test Power: A Product Team's Guide to CUPED and Variance Reduction

Learn how CUPED significantly reduces experiment variance, enabling product teams to detect smaller effects with smaller sample sizes and greater statistical power, accelerating data-driven decisions.

Quantifying and Mitigating User Interference & Network Effects for Robust A/B Testing

Accurately measuring experiment outcomes requires understanding and addressing user interference and network effects. This guide outlines practical methodologies for identifying, mitigating, and quantifying these phenomena using robust effect-size benchmarks.

How to control for multiple comparisons in A/B/n testing

How to control for multiple comparisons in A/B/n testing

Defining Your OEC: A Practical Metric-Framework for Robust Experimental Design and Primary Metrics

Establishing a clear Overall Evaluation Criterion (OEC) is fundamental for effective experimental design, ensuring that experiments drive meaningful business outcomes. This guide outlines a precise methodology for PMs, growth teams, and analysts to define a singular, actionable OEC.

Handling Outliers in A/B Testing: A Practical Guide for Revenue & Engagement Data

Outliers can severely skew A/B test results, leading to misinformed product and business decisions. This guide provides practical methodologies like winsorization and truncation to responsibly manage extreme values in revenue and engagement metrics.

Non-Inferiority A/B Testing: A Practical Guide for Product & Growth Teams

Non-inferiority A/B tests are crucial for validating new features without compromising key metrics. This guide outlines the methodology to design, run, and interpret these guardrail tests effectively.

Defining funnel steps consistently across teams

Defining funnel steps consistently across teams

Detecting Novelty Effects in Long-Running Experiments: A Practical Methodology Guide

Understand how to identify and address transient novelty effects in long-running A/B tests to prevent misleading conclusions. This guide outlines practical steps for PMs, growth teams, and analysts to ensure robust experiment outcomes.

Power Analysis Primer for Product Experiments: A Practical AB Test Methodology

Understand how to conduct a robust power analysis for product AB tests, ensuring your experiments are adequately sized to detect meaningful effects. This guide outlines key inputs, process, and how to leverage StatFacts benchmarks for more precise experimental design.

Effective Benchmark Segmentation: Device and Traffic Source Strategies

Accurate performance evaluation requires segmenting benchmarks by critical dimensions like device type and traffic source. This guide details a methodical approach to ensure your comparative insights are relevant and actionable.

Documenting external priors in experiment briefs

Use industry benchmarks and past tests to set realistic effect-size expectations before you launch an A/B test.

Selecting Robust Guardrail Metrics for A/B Tests: A Practical Methodology

Guardrail metrics are crucial for preventing unintended negative consequences in A/B tests. This guide provides a practical, step-by-step methodology for selecting, defining, and monitoring these essential safeguards to ensure responsible experimentation.

Meta-Analysis for Product Teams: Understanding Its Power and Limits

Meta-analysis can illuminate general trends and average effects across studies, offering valuable benchmarks for product decisions. However, product teams must understand its inherent limitations regarding specificity, context, and applicability to their unique user base and product features.

Detecting Sample Ratio Mismatch (SRM) in Experiments: A Methodological Guide

Sample Ratio Mismatch (SRM) can invalidate experiment results, leading to flawed decisions. This guide outlines practical steps for its detection and interpretation, ensuring the integrity of your A/B tests and observational studies.

Precise Benchmarking: Adjusting for Seasonality and Campaign Context

Static benchmarks rarely reflect dynamic business realities. Learn how to meticulously adjust your performance benchmarks to accurately account for the predictable shifts of seasonality and the targeted impacts of marketing or product campaigns, ensuring your evaluations remain relevant and actionable.

How to Read and Interpret Benchmarks Correctly

A short guide to effect labels, confidence tags, and sample context—before you paste a number into Slack.

Relative vs Absolute: Understanding Difference, Change & Risk

Why +10% and +10 points are not the same—and how to avoid embarrassing math in your next deck.

Confidence Levels Explained: Understanding Statistical Evidence

What meta-analysis, A/B test, study, and estimate mean on StatFacts—and how much weight to give each.

When not to use a benchmark

Five situations where copying a StatFacts number will mislead your team—and what to do instead.

How to Cite Sources and Statistics in a Slide Deck

Copy-paste templates for Slack, Notion, and slides—so your numbers look credible, not copied.

Planning an A/B test from a benchmark

Turn a StatFacts range into a test hypothesis, success metric, and realistic minimum detectable effect.