// Deterministic go/no-go for every pull request

Know if a pull request is safe to ship before you merge.

robustiq reads the diff and your CI history, scores risk and confidence, and flags anything no test has verified. Then it checks the result against your release policy, so you get a go/no-go you can explain today and still audit later.

Works with the git and CI you already run. No new tests to write.

Illustrative example - not customer data.

Works with the tools you already run

  • GitHub
  • GitLab
  • GitHub Actions
  • Jenkins
  • JUnit XML
  • Slack
  • Claude Code · MCP
  • Cursor · MCP

// The problem

Passing tests don't tell you what's safe to ship.

CI tells you which tests passed. It doesn't tell you whether any of them touch what changed, whether last night's red build was real or just flaky, or how this diff compares with the last fifty. So the release call gets made in a chat thread, and a month later nobody can say why it went out.

01

Flaky tests teach people to ignore red.

When a test fails one run in ten, people just hit re-run, and real failures slip through along with the noise.

02

Risk lives in the diff, not in averages.

Repo-wide coverage and pass rates say little about the 40 lines this pull request changes.

03

Release calls leave no trail.

When something breaks, there's no record of who decided, or what they knew and checked at the time.

04

More code, same reviewers.

AI coding agents open pull requests faster than people can review them. The bottleneck has moved from writing code to deciding what's safe to merge.

// How it works

From diff to decision in four steps.

  1. 01

    Point it at a change.

    Any base..head diff: a pull request, a branch or a commit range. Run it in CI on each pull request, or ask from an AI agent over MCP.

    base main .. head feature/payment-retry
  2. 02

    It scores the change.

    A deterministic pipeline computes risk, confidence and evidence coverage. It also names the top risk drivers and lists the risks nothing has verified.

    risk62conf.71cov.58%
  3. 03

    Your policy decides.

    Each gate compares an actual value with the threshold your team set. A HOLD or NO-GO always names the gate that failed.

    coverage ≥ 70%58%FAIL
  4. 04

    The decision is kept.

    Each decision is stored in a durable log and shows up on the Trends dashboard. You can also package it as an evidence bundle.

    #482 · HOLD · recorded · bundle ready

Illustrative example - not customer data.

// Anatomy of a decision

Every verdict comes with its reasons.

This is what robustiq returns for a single pull request. Each number traces back to inputs you can inspect, so nothing is hidden behind a score.

robustiq report --base main --head feature/payment-retry--format full
  1. 1HOLDPR #482 · 2 of 4 gates failed
  2. 2
    Risk62/100
    Confidence71/100
  3. 3
    Evidence coverage · changed lines and behaviours58%
    coveredpartialunverified
  4. 4
    Top risk driverspayments/retry.ts - new retry path, no test exercises itcheckout.spec › retries once on 503 - flake score 64
  5. 5
    Unverified risks · 3Retry fires twice when the gateway times outIdempotency key not reused on the second attemptRefund path after a partial capture
  6. 6
    risk ≤ 7062PASSconfidence ≥ 6071PASSevidence coverage ≥ 70%58%FAILunverified risks ≤ 13FAIL
  1. 01

    Decision

    GO, HOLD or NO-GO for each pull request.

  2. 02

    Risk and confidence

    How risky the change is, and how much evidence you have about it.

  3. 03

    Evidence coverage

    How much of what changed is covered by evidence you trust.

  4. 04

    Top risk drivers

    The specific files and behaviours that push the risk up.

  5. 05

    Unverified risks

    What nothing has checked yet. These are the questions to ask before the merge, not after an incident.

  6. 06

    Policy gates

    Actual vs. threshold for every gate, so nobody has to argue about why a pull request was held.

  7. Illustrative example - not customer data.

// What you get

Everything behind the decision, on record.

Flaky tests

Know which red builds to trust.

Each test gets a 0-100 flake score based on your CI history. Start by fixing the worst offenders.

  1. checkout.spec › retries once on 50364
  2. auth.spec › refresh token rotates41
  3. search.spec › debounces input23
  4. cart.spec › applies coupon7
Policy gates

Your release rules, as thresholds you can read.

Set them once, then see actual vs. threshold on every pull request.

risk ≤ 70confidence ≥ 60evidence coverage ≥ 70%unverified risks ≤ 1+ add gate
Decision log

Every decision, in one log.

A durable log by branch and pull request, with a Trends view of how risk and confidence move over time.

  • #482 · feature/payment-retryHOLD
  • #481 · fix/session-timeoutGO
  • #479 · chore/deps-bumpGO
Evidence

Audit-ready by default.

Package each decision with the evidence behind it for change approvals and post-incident reviews.

evidence/pr-482/├─ decision.json HOLD├─ gates.json 2/4 failed├─ risk-drivers.json├─ unverified.json 3 items└─ inputs.sha256

File layout illustrative.

Illustrative example - not customer data.

// For AI coding agents

Let your coding agent check its work before it asks for review.

robustiq ships an MCP server, so any MCP-compatible agent can ask what your CI asks and get the same deterministic answer.

  • robustiq_reportRelease decision for a base..head diff, as a summary or the full gate table.
  • robustiq_flaky_testsFlake scores from your CI history, filterable by minimum score.
  • robustiq_decision_trendRecent decisions from the decision log, by branch or pull request.

Read-only by design: MCP calls never record decisions, write evidence bundles or send notifications, so your audit trail stays under your CI's control.

Claude Code, Cursor, any MCP client
agent session · illustrative
you › Is my branch safe to merge into main?
agent → robustiq_report(base: "main", head: "HEAD")
← HOLD · evidence coverage 58% is below your 70% gate3 unverified risks in payments/retry.ts
agent › Want me to add tests for those three paths first?
you ›

// How it decides

Deterministic, not a black box.

Run the same diff, history and policy twice and you get the same decision. Each score traces back to inputs you can inspect, which is what you want from anything deciding what reaches production.

1decision per pull request
0-100flake score per test, from your CI history
actual vs. thresholdfor every gate, on every report
same input → same decisionevery time

What robustiq reads: your git repository and CI test results.

Security

// Built in the open

We gate robustiq with robustiq.

Every pull request to our own codebase gets a decision before it merges.

// FAQ

Questions engineers ask first.

Something else? Ask us on a call.

Is robustiq a test automation tool?

No. robustiq doesn't write or run your tests. It reads the evidence your tests and CI already produce and turns it into a release decision.

Is the decision made by an AI model?

No. The pipeline is deterministic: the same diff, history and policy always produce the same decision, and every score traces back to its inputs.

What does robustiq need access to?

Your git repository, to compare two refs, and your CI test history.

How does it handle flaky tests?

Each test gets a 0-100 flake score from your CI history, so you can see which failures are noise and fix the worst offenders first.

Can we set our own release policy?

Yes. Gates are thresholds your team chooses, and every full report shows each gate's actual value next to its threshold.

Can AI coding agents use it?

Yes, through robustiq's MCP server. Agents get the same decision your CI gets, and the MCP tools are read-only.

Make your next release call on evidence instead of a hunch.