// Product

From diff to decision, with a record of why.

robustiq compares a base and a head ref, reads your CI test history and scores the change for risk, confidence and evidence coverage. Then it checks the change against your policy gates. You get one decision that names the failed gates and the risks nothing has verified, and you can record it to a durable log to query later.

robustiq · one change, end to endPR #482
  1. 1
    Compare two refsbase main .. head feature/payment-retry
  2. 2
    Read CI test historycheckout.spec › retries once on 503 · flake score 64
  3. 3
    Score the changerisk 62 · confidence 71 · coverage 58%
  4. 4
    Name what nothing verified2 top risk drivers · 3 unverified risks
  5. 5
    Check your policy gates4 gates · 2 failed: evidence coverage, unverified risks
  6. 6
    DecideHOLD
  7. 7
    Keep the recorddecision log · evidence bundle · notification
Illustrative example - not customer data.

// The full report

The verdict and the numbers behind it, on one page.

Ask for the full format and the summary gets a policy gate table with each gate's actual value, threshold and result. Here is the full report for one pull request.

robustiq_report(base: "main", head: "feature/payment-retry", format: "full")Illustrative example
PR #482 · Add retry to payment capturefeature/payment-retry → main
2 of 4 gates failedHOLD
Riskgate ≤ 70
62/100
Confidencegate ≥ 60
71/100
Evidence coveragegate ≥ 70%
58%
Top risk drivers · 2
payments/retry.ts - new retry path, no test exercises it
checkout.spec › retries once on 503 - flaky, flake score 64
Unverified risks · 3
  1. 01Retry fires twice when the gateway times out
  2. 02Idempotency key not reused on the second attempt
  3. 03Refund path after a partial capture
Policy gates · actual vs. threshold
Policy gateactualthresholdresult
risk ≤ 706270PASS
confidence ≥ 607160PASS
evidence coverage ≥ 70%58%70%FAIL
unverified risks ≤ 131FAIL

Both score gates pass. Evidence coverage and unverified risks hold the pull request back, and the report says so.

Illustrative example - not customer data.

// Risk and confidence

Two scores, because risky and unknown are different problems.

A change can be risky and well tested, or harmless and barely tested. robustiq keeps the two numbers apart, so you know whether to slow down or to gather more evidence.

Risk · y-axis
0-100

How risky the change is.

What the diff touches and how much it changes. The report names the top risk drivers: the files and behaviours pushing the number up.

#48262
Confidence · x-axis
0-100

How much evidence you have about it.

How much of what changed is backed by tests that exercise it and a CI history you can trust. Low confidence means nobody has checked the code yet, not that it's bad.

#48271
risk × confidence · 12 recent pull requests · base mainIllustrative

#482 sits in the ship corner on scores alone but still gets a HOLD. Evidence coverage and unverified risks failed their gates. The scores feed into the decision, but your gates make the call.

Illustrative example - not customer data. Dashed lines are the example gates: risk ≤ 70, confidence ≥ 60.

// Unverified risks

What nothing has checked, named before the merge.

Evidence coverage tells you how much of a change is covered. Unverified risks tell you what isn't, one behaviour at a time, so a reviewer knows what to ask and a coding agent knows which test to write.

  • Listed, not just counted.Each item names one behaviour the diff changes that nothing has checked.
  • Gate them if you want to.A gate such as unverified risks ≤ 1 holds the change until the list is short enough. PR #482 has three.
  • Readable by agents.robustiq_report returns the same list over MCP, so an agent can close the gaps before it asks for review.
unverified risks · PR #482 · 3gate ≤ 1 · FAIL
  1. Retry fires twice when the gateway times outpayments/retry.ts › captureWithRetry()
    Suggested check
    Time out the gateway after the first attempt; assert exactly one capture.
  2. Idempotency key not reused on the second attemptpayments/retry.ts › buildRetryRequest()
    Suggested check
    Assert the second attempt sends the first attempt's idempotency key.
  3. Refund path after a partial capturepayments/refunds.ts › refundCapture()
    Suggested check
    Refund an order whose capture was partial and retried; assert the refunded amount.
Illustrative example - not customer data.

// Policy gates

You set the thresholds. robustiq measures.

A gate pairs one measured value with a threshold your team chooses. The full report puts each gate's actual value next to its threshold, so a held pull request comes with the number that held it.

release policy · 4 gates Illustrative editor
Illustrative release policy: four gates with comparator, threshold, scope and the last result on pull request 482
GateRuleThresholdScopeLast · #482
risk≤70branch: main62 PASS
confidence≥60branch: main71 PASS
evidence coverage≥70%branch: main58% FAIL
unverified risks≤1branch: main3 FAIL
add gate
Illustrative example - not customer data.
  • 01

    Actual next to threshold.

    The report prints each gate's actual value and threshold, plus PASS or FAIL.

  • 02

    Failures are named.

    A HOLD or NO-GO always names the gate that failed. Nobody has to reverse-engineer a score.

  • 03

    One policy for CI and agents.

    An agent asking over MCP is checked against the same gates as your pipeline.

// Decision log & Trends

Every decision on record, and which way they're trending.

Recorded decisions go into a durable log you can filter by branch or pull request. The dashboard's Trends tab charts risk and confidence across that log, so you can spot a slow slide before it turns into an incident.

robustiq · dashboardTrendsIllustrative example
Decision logbranch: main · 5 most recent
Illustrative decision log for branch main, five most recent decisions
PRHead branchDecisionRiskConf.Recorded
#482feature/payment-retryHOLD6271Oct 1 · 14:32
#481fix/session-timeoutGO2482Oct 1 · 11:05
#480feature/ledger-migrationNO-GO8138Sep 30 · 17:48
#479chore/deps-bumpGO1888Sep 30 · 09:12
#478refactor/cart-totalsHOLD4756Sep 29 · 16:20
Illustrative example - not customer data.Agents read the same log: robustiq_decision_trend by branch or pull request.

// Evidence bundles

Audit-ready: the decision and its evidence, in one bundle.

For a recorded decision, robustiq can write an evidence bundle: the verdict, the gate results, the risk drivers and the unverified risks it was based on. You attach it to a change approval, and open it again in a post-incident review.

Change approval

Attach the bundle to the change request. The approver sees the evidence the gates saw, not a screenshot of a green build.

Post-incident review

Open the bundle for the release in question: what was known at merge time, which gates passed, and what nothing had verified.

evidence bundle · PR #482Illustrative
  • evidence/pr-482/
  • ├─ decision.jsonHOLD
  • ├─ gates.json4 gates · 2 failed
  • ├─ risk-drivers.json2 drivers
  • ├─ unverified.json3 items
  • └─ inputs.sha256base, head, CI history
decision.jsonPreview
{
  "pull_request": 482,
  "base": "main",
  "head": "feature/payment-retry",
  "decision": "HOLD",
  "risk": 62,
  "confidence": 71,
  "evidence_coverage": 0.58,
  "failed_gates": ["evidence_coverage", "unverified_risks"],
  "unverified_risks": 3
}
Illustrative example - not customer data. File layout and fields illustrative.

// Notifications

The right people hear about a HOLD.

robustiq can send a notification when a decision is recorded, so a held pull request doesn't sit unnoticed in a CI log.

robustiqdecision recorded · just now
PR #482 · Add retry to payment capture
HOLDfeature/payment-retry → main
evidence coverage 58% < 70% · unverified risks 3 > 1
Illustrative example - not customer data. Channel-neutral preview.

// Deployment & data

What robustiq reads and what it keeps.

Each decision is computed from two sources (your git repository and your CI test history) and checked against the policy you set. Here's what that means for your code and your data.

Reads

Your repository and your CI history.

Git repository
to compare a base ref with a head ref
CI test history
ingested results for evidence and flake scores
Your policy
the gates and thresholds your team sets
Keeps

What it keeps.

Decision log and evidence bundles
for the decisions you record
MCP calls
read-only: no records, bundles or notifications

See the full report for one of your own pull requests.