Decision
GO, HOLD or NO-GO for each pull request.
// Deterministic go/no-go for every pull request
robustiq reads the diff and your CI history, scores risk and confidence, and flags anything no test has verified. Then it checks the result against your release policy, so you get a go/no-go you can explain today and still audit later.
Works with the git and CI you already run. No new tests to write.
| Policy gate | actual | result |
|---|---|---|
| risk ≤ 70 | 62 | PASS |
| confidence ≥ 60 | 71 | PASS |
| evidence coverage ≥ 70% | 58% | FAIL |
| unverified risks ≤ 1 | 3 | FAIL |
Works with the tools you already run
// The problem
CI tells you which tests passed. It doesn't tell you whether any of them touch what changed, whether last night's red build was real or just flaky, or how this diff compares with the last fifty. So the release call gets made in a chat thread, and a month later nobody can say why it went out.
When a test fails one run in ten, people just hit re-run, and real failures slip through along with the noise.
Repo-wide coverage and pass rates say little about the 40 lines this pull request changes.
When something breaks, there's no record of who decided, or what they knew and checked at the time.
AI coding agents open pull requests faster than people can review them. The bottleneck has moved from writing code to deciding what's safe to merge.
// How it works
Any base..head diff: a pull request, a branch or a commit range. Run it in CI on each pull request, or ask from an AI agent over MCP.
A deterministic pipeline computes risk, confidence and evidence coverage. It also names the top risk drivers and lists the risks nothing has verified.
Each gate compares an actual value with the threshold your team set. A HOLD or NO-GO always names the gate that failed.
Each decision is stored in a durable log and shows up on the Trends dashboard. You can also package it as an evidence bundle.
Illustrative example - not customer data.
// Anatomy of a decision
This is what robustiq returns for a single pull request. Each number traces back to inputs you can inspect, so nothing is hidden behind a score.
GO, HOLD or NO-GO for each pull request.
How risky the change is, and how much evidence you have about it.
How much of what changed is covered by evidence you trust.
The specific files and behaviours that push the risk up.
What nothing has checked yet. These are the questions to ask before the merge, not after an incident.
Actual vs. threshold for every gate, so nobody has to argue about why a pull request was held.
// What you get
Each test gets a 0-100 flake score based on your CI history. Start by fixing the worst offenders.
Set them once, then see actual vs. threshold on every pull request.
A durable log by branch and pull request, with a Trends view of how risk and confidence move over time.
Package each decision with the evidence behind it for change approvals and post-incident reviews.
File layout illustrative.
Illustrative example - not customer data.
// For AI coding agents
robustiq ships an MCP server, so any MCP-compatible agent can ask what your CI asks and get the same deterministic answer.
Read-only by design: MCP calls never record decisions, write evidence bundles or send notifications, so your audit trail stays under your CI's control.
// How it decides
Run the same diff, history and policy twice and you get the same decision. Each score traces back to inputs you can inspect, which is what you want from anything deciding what reaches production.
What robustiq reads: your git repository and CI test results.
Security// Built in the open
Every pull request to our own codebase gets a decision before it merges.
No. robustiq doesn't write or run your tests. It reads the evidence your tests and CI already produce and turns it into a release decision.
No. The pipeline is deterministic: the same diff, history and policy always produce the same decision, and every score traces back to its inputs.
Your git repository, to compare two refs, and your CI test history.
Each test gets a 0-100 flake score from your CI history, so you can see which failures are noise and fix the worst offenders first.
Yes. Gates are thresholds your team chooses, and every full report shows each gate's actual value next to its threshold.
Yes, through robustiq's MCP server. Agents get the same decision your CI gets, and the MCP tools are read-only.