CarKey CasePilotBuild Week Demo

Synthetic judge workspace

Evaluation-only workflow using fictional data. It is separate from the public customer journey.

Reproducible technical evidence

Safety and quality are tested as executable contracts.

A local, deterministic evaluation suite measures CasePilot's structured analysis, refusal behavior, privacy masking, authorization gates, bilingual consistency, and verified-fact grounding — with no API key or network request.

All current checks passedcasepilot-deterministic-contract-v1DeterministicDemoProvider

73

Contract checks

73

Passed

14

Analysis fixtures

0

Network calls

Evaluation suite

Measured reliability dimensions

100%

Schema-valid structured output14/14
Expected missing fields and risks14/14
Sensitive-request refusal14/14
Deterministic replay14/14
English / zh-TW behavioral parity3/3
PII masking4/4
Authorization and dispatch gates6/6
Grounded, human-reviewed Growth Pack4/4

This is a reproducible contract benchmark, not a claim of live-model or real-world accuracy.

How the evidence works

1. Synthetic golden set

English and Traditional Chinese intake samples, ambiguous requests, six sensitive-request categories, and four PII patterns define expected behavior without customer data.

2. Deterministic graders

Exact schema, classification, replay, state-transition, privacy-boundary, bilingual, grounding, and publication checks produce the score shown above.

3. Safe expansion path

Owner-approved work material can contribute generalized facts to new synthetic fixtures. Raw customer media and identifiers stay outside this suite.

What this evidence does not claim

The public judge experience runs DeterministicDemoProvider. These results do not claim live GPT-5.6 runtime accuracy, production conversion lift, real-customer validation, or permission to publish generated content.

Run the guided workflowReview safety boundaries