Experiments

Experiments are how you test beliefs against reality. Every website experiment answers one question: does Apex change the page, or does your code? Assignment is sticky either way. First paint is a separate problem.

Who applies the variant

  • Apex changes the page — the tracking snippet swaps text or DOM that is already on the page. No deploy. No npm package. The snippet hides the page for a moment so the visitor does not see that swap.
  • Your code changes the page — you call useApexVariant and render the arm. It can change layout, components, or server output. It starts collecting once that code is deployed.

Tip

If the snippet is already on a marketing site, let Apex change the page. If you are in a Next or React app and the variant is real UI, your code applies it.

Experiment Structure

Every experiment has these core fields:

FieldDescription
nameA descriptive name (e.g. "Pricing page social proof")
targetUrlThe page where the experiment runs
trafficSplitPercentage of visitors included (e.g. 100 for all traffic)
variantsAt least two. Website experiments usually start with control and variant_b. Journey and letter experiments can have more.
statusOne of: draft, running, paused, completed
hypothesisThe testable claim
conversion / metricThe Conversion that defines success

Traffic is split evenly between control and variant_b by default (50/50). You can adjust this, but even splits give you the fastest path to statistical significance.

How Assignment Works

Apex uses MurmurHash3 for deterministic visitor bucketing:

  1. Each visitor gets a persistent anonymous ID (apex_vid)
  2. The hash of visitorId + experimentId produces a number between 0 and 99
  3. That number picks control or the variant
  4. The same visitor always gets the same arm. That is sticky assignment. It is not first paint.

Snippet experiments assign in the browser and apply the change before the page is shown.

Code experiments must assign on the document request (same rules as GET /api/experiments/{id}/assign) and pass that arm into useApexVariant(id, { initialVariant }) or <ApexProvider assignments={{ [id]: variant }}>. If you use Next.js, npm install @apex-inc/next @apex-inc/react, then apexVisitorMiddleware, assignExperiments, <ApexProvider assignments>. The hook alone is not flicker-free. Without the server arm, the first HTML is control; after /assign returns, React swaps. That swap is the flicker. A client-only app cannot promise no flicker.

Creating the experiment in the dashboard or through MCP does not add that wiring.

Info

Sticky assignment means they do not bounce between arms on reload. First HTML is flicker-free only when it is already the assigned arm. See Running Experiments.

Running an Experiment

Create the experiment

From the dashboard, click New Experiment. Give it a name, choose the target URL, and select the variant type.

Configure the variant

If Apex changes the page, set the selector and the new copy in the visual editor. If your code changes the page, implement the arm with useApexVariant and pass a server assign so first HTML is already that arm.

Pick the conversion

Choose an existing conversion or create a new one. That is the metric Apex uses to pick a winner.

Link a belief (optional)

Connect the experiment to a belief to automatically update confidence when results come in.

Activate

Set the status to running. Snippet experiments start assigning immediately. Code experiments should have the server assign shipped first; otherwise visitors can see a flash of control.

Results and Statistical Confidence

As visitors flow through the experiment, Apex tracks:

  • Visitor count per variant
  • Conversion count and conversion rate per variant
  • Relative lift (how much better or worse the variant performs vs control)
  • Statistical confidence (how likely the observed difference is real, not noise)

Apex picks the right estimator for your metric. Results are considered significant at 95% confidence by default. You'll see the confidence percentage climb as more data comes in.

Warning

Don't call experiments early. Statistical significance requires enough data — calling a winner at 80% confidence means there's a 1-in-5 chance you're wrong. Wait for 95%.

The measurement contract

There is one measurement contract with the right estimator behind it for each metric type, so a conversion is always scored honestly:

Metric typeEstimatorNotes
rate (did it happen?)Two-proportion z-testConversion rate vs control
countz-testPer-subject counts
revenue / duration (how much?)Welch's t-test on per-subject valuesUnequal-variance, with a 95% CI on the difference of means
Adaptive comm / journey armsBayesian (Beta-Binomial, probability-to-be-best)Powers Thompson sampling

Whichever engine runs, the output is the same single field — the probability this winner is actually best — so the readiness gate and calibration read it identically.

Continuous metrics are handled honestly. Revenue and duration are heavy-tailed, so Apex winsorizes the extremes (a few whales can't decide the test) and sums value per subject, not per event. Revenue conversions can also be net (purchase value minus refunds).

Power before you launch. When you pick a conversion, Apex estimates how long the test will take to reach significance from your live event volume and warns when a design is underpowered ("~N weeks; reduce variants or raise the minimum detectable effect"). Testing more variants or more metrics tightens the bar automatically (a multiple-comparisons correction that scales with the number of arms and guardrails).

Value maturation. Revenue and other late-settling outcomes don't become decisive the moment they're statistically ahead — Apex holds the call through a maturation window (default 30 days) so a winner isn't crowned on revenue that could still be refunded. While it waits, the experiment shows a distinct maturing state.

Frozen at launch. When an experiment starts running, its primary metric and guardrails are frozen. Editing them requires forking a new experiment — this prevents moving the goalposts mid-test (HARKing).

Guardrails (protected metrics)

Every experiment is protected by default. Apex watches a set of business "vital signs" and blocks a variant from winning if it causes harm — you don't configure anything.

  • What's watched is derived from your conversion model plus universal harms, and only attached when the event is actually firing: revenue (must not drop), refunds and errors/crashes (must not rise), plus model-specific signals (activation for PLG, demo requests for sales-led, supply/demand signups for marketplaces). Add extra guardrails under Advanced when a test has a specific risk.
  • How a breach is judged. Guardrails are checked continuously, so Apex uses an always-valid (sequential) test — not a fixed-horizon one — so watching every poll doesn't inflate false alarms. A breach is only critical when the harm is past your safe-range margin AND statistically confirmed AND past a sample floor. Revenue/refund guardrails also respect the maturation window. Anything less shows as At risk, not a block.
  • SRM. Apex also checks that traffic is splitting the way you intended (sample-ratio-mismatch). A broken split means the results are invalid, so Apex flags the experiment rather than trusting the numbers.
  • What happens on a breach. Promotion is blocked and you get a notification. If you turn on auto-pause (Settings or the "Protect your experiments" setup step), Apex pauses the experiment automatically; otherwise it blocks + alerts and re-alerts daily until you decide.

Muting the notification never disables the protection — the block and (optional) auto-pause always apply.

Code experiments (Next.js)

For a Next app, assign on the document request and pass the arm in. Do not install @apex-inc/sdk for this. That package tracks events. It does not assign the first HTML.

import { assignExperiments } from "@apex-inc/next";
import { ApexProvider, useApexVariant } from "@apex-inc/react";

const { assignments } = await assignExperiments(["exp_…"], { visitorId });

<ApexProvider assignments={assignments}>
  <Hero />
</ApexProvider>

const variant = useApexVariant("exp_…", { initialVariant: assignments["exp_…"] });

assignExperiments uses the same bucketing as the snippet, so a visitor stays on the same arm. The hook still re-checks /assign after paint; a sticky assignment does not change what they first see.

On-device screenshots (mobile)

Apex auto-captures public web URLs server-side when you create an experiment, and the agent CLI (via attach_experiment_asset) covers localhost. But servers can't reach authed in-app mobile screens — so the Capacitor plugin can capture them on-device.

On the screen that renders the variant, call captureVariantScreenshot() keyed to the resolved variant:

const variant = useApexVariant(experimentId, { initialVariant }); // or: const { variant } = await Apex.getVariant({ experimentId })
useEffect(() => {
  if (variant) {
    Apex.captureVariantScreenshot({ experimentId, variantKey: variant });
  }
}, [variant]);

Info

captureVariantScreenshot() is debug-gated — it only fires when the plugin is initialized with Apex.initialize({ ..., debug: true }). In production it no-ops ({ captured: false, reason: "debug_disabled" }) and never screenshots real end users. It's a QA/review tool.

The captured shot lands on the dashboard experiment card (cover) and the detail Variant previews gallery, right alongside web and agent captures.

Grounded experiments (brand truth)

When an AI agent writes experiment copy, it should never invent facts. Apex ships four short, merchant-owned guardrail files you drop into your repo so the agent writes variants against your truth:

FileWhat it holds
.apex/brand-truth.mdThe only facts experiments may claim (pricing, shipping, perks). If it's not here, the agent doesn't say it.
.apex/brand-voice.mdHow copy should sound.
.apex/lexicon.mdPreferred / banned / qualifier-required words.
.apex/experiment-rules.mdDesign + implementation discipline.

A fifth file, .cursor/rules/apex-experiment-guardrails.mdc, is an always-on rule that tells the agent to read the four files before writing any experiment.

Setting them up

In the dashboard, open Set up Apex → Intelligence → "Set up experiment guardrails." The recommended (Cursor) lane scaffolds all five files into your repo through your connected agent; a non-Cursor lane lets you download them from /shipped-guardrails (with a SHA-256 manifest to verify integrity). The templates ship blank — you fill in the facts. Apex never guesses your facts for you.

How enforcement works

There's no server-side fact-checker. Enforcement is by use: create_experiment and the new-experiment prompt instruct the agent to read .apex/brand-truth.md and cite the line each claim traces to before proposing copy. If a claim isn't backed there, the agent asks instead of inventing.

Warning

Why this matters: an agent once wrote "Free shipping on every order" for a store that charges $8 under $100 — a plausible-sounding claim the product couldn't back, which would have invalidated the test. brand-truth.md is what keeps that from shipping.

Where an experiment runs: data source + randomization unit

Every experiment belongs to a data source (a property: your website, your iOS app, your Android app, your backend) — the same registry that classifies where your events come from. This is what the experiment runs in, and it's separate from surface (the delivery channel: a DOM swap vs a native payload). A Capacitor app is one property, so all of its experiments group under the one mobile app — never split across "Website" and "Mobile app".

Each experiment also has a randomization unit — what it buckets on:

  • visitor (default) — one anonymous device/browser.
  • person — one stitched human; all their devices get the same variant (requires identify()).
  • account / org — everyone in a B2B account gets the same variant.

The randomization unit is also the unit results are counted on, so pick account for account-level tests or the math treats four devices in one account as four independent draws.

Counting one user across web, iOS, and Android (golden path)

Apex can only treat one human as one subject if you give them one id. The golden path (same as Statsig/PostHog):

  1. Mint one Apex visitor id and pass it everywhere. The SDKs persist it under apex_vid; for a web→app transition, carry it across and call Apex.setVisitorId(webVid) so the app reuses the site's id.
  2. Call identify({ email }) on every platform after login. Apex links the per-device ids to one person on verified email — it never merges on device signals alone.
  3. Server-side events: pass the browser's apex_vid as the visitor id so server events join the same person.

Assignment is sticky: once a subject is bucketed, that variant is locked — re-evaluations never re-bucket them.

Connecting to Predictions

Before running an experiment, log a prediction — what you think will happen. This builds your team's calibration score and makes experiment results more actionable.

After the experiment completes, Apex compares your prediction against actual results to calculate an accuracy score. Over time, this feedback loop makes your team better at anticipating outcomes.

Lifecycle

StatusWhat's happening
draftExperiment is configured but not live. No visitors are assigned.
runningVisitors are being assigned and tracked. Results update in real-time.
pausedAssignment stops. Existing data is preserved. Can be resumed.
completedExperiment is finished. Results are final. Belief confidence is updated.