Executable evidence

Proof Lab

An interactive engineering lab where visitors can execute field-level evaluation, attack bounded agent plans, disable architecture safeguards, and profile their own structured data without uploading it.

Executable extraction evaluation

Edit synthetic candidate output and run deterministic field-level precision, recall, F1, exact-match, missing-field, and unexpected-field checks against an explicit release gate.

Data Analyst Agent control room

Edit adversarial tool calls and run a browser contract mirror of the public Data Analyst Agent repository’s 33-tool allowlist, Pydantic argument bounds, dataset-column grounding, and deterministic executor boundary. The reviewed repository defines 486 default tests across 51 modules and 21 end-to-end tool workflows.

Architecture Chaos Mode

Inject bounded failure scenarios into published project architectures, disable individual safeguards, and trace detection, containment, recovery, and broken invariants. The result is explicitly labelled as simulation.

Private data challenge

Profile CSV, TSV, or JSON locally in the browser. File contents are not uploaded; the lab produces concrete row, column, completeness, duplicate, type, and validation-contract results.

Versioned proof receipts

Issue and verify a shared receipt for evaluation, chaos, data, policy, and claim evidence. The verifier recomputes canonical SHA-256 integrity, checks the currently published release manifest, and keeps build provenance separate from claim truth.

Read-only portfolio MCP

Connect an external agent to six typed, read-only tools for public evidence search, architecture inspection, project comparison, claim challenge, provenance, and deterministic evaluation. The server cannot inspect repositories, call a model, run arbitrary SQL, or mutate state.

Inspect the machine-readable MCP capability manifest.

Review the AI trust boundaries or inspect claim-level evidence.