Behavior passed ≠ prompt workedFive transparent trials

Did your prompt work—or did GPT‑5.6 save it?

Paste the instructions that govern your agent. TestForge separates what the model did from what your prompt actually supports—then shows the evidence and smallest useful redlines.

Enter the trial bay
01 / Supply the control surface

What tells this agent how to behave?

Start with a labeled example or paste your own control surface. Live runs use five independent subject trials and a separate evidence judgment.

Evidence boundary

A behavioral trial is evidence from named cases, not a safety certification or guarantee of deployment behavior.

02 / Observe behavior under pressure

The trial record will resolve here.

Each case gets two separate judgments: what the agent did, and whether its instructions deserve the credit.

Select an example and run the suite. No overall score will appear—only case-level evidence.

Method / TestForge inside

Behavior first. Receipts attached.

Agent Trials is a public TestForge surface—not a separate certification product and not a universal benchmark.

01

Isolated pressure

Each trial changes one consequential condition and keeps the behavioral target legible.

02

Two honest verdicts

Behavior says what the model did. Control says whether the submitted instructions deserve the credit.

03

Named evidence

Every verdict cites the response, observed decisions, criterion, and instruction clause that earned it.

04

Minimal repair

Redlines target the exposed instruction gap. A repaired case earns a retest, never a fictional license.