Executable evidence
Proof Lab
An interactive engineering lab where visitors can execute field-level evaluation, attack bounded agent plans, disable architecture safeguards, and profile their own structured data without uploading it.
Executable extraction evaluation
Edit synthetic candidate output and run deterministic field-level precision, recall, F1, exact-match, missing-field, and unexpected-field checks against an explicit release gate.
Data Analyst Agent control room
Edit adversarial tool calls and run a browser contract mirror of the public Data Analyst Agent repository’s 33-tool allowlist, Pydantic argument bounds, dataset-column grounding, and deterministic executor boundary. The reviewed repository defines 486 default tests across 51 modules and 21 end-to-end tool workflows.
Architecture Chaos Mode
Inject bounded failure scenarios into published project architectures, disable individual safeguards, and trace detection, containment, recovery, and broken invariants. The result is explicitly labelled as simulation.
Private data challenge
Profile CSV, TSV, or JSON locally in the browser. File contents are not uploaded; the lab produces concrete row, column, completeness, duplicate, type, and validation-contract results.
Versioned proof receipts
Issue and verify a shared receipt for evaluation, chaos, data, policy, and claim evidence. The verifier recomputes canonical SHA-256 integrity, checks the currently published release manifest, and keeps build provenance separate from claim truth.
Read-only portfolio MCP
Connect an external agent to six typed, read-only tools for public evidence search, architecture inspection, project comparison, claim challenge, provenance, and deterministic evaluation. The server cannot inspect repositories, call a model, run arbitrary SQL, or mutate state.
Review the AI trust boundaries or inspect claim-level evidence.