73
Contract checks
Synthetic judge workspace
Evaluation-only workflow using fictional data. It is separate from the public customer journey.
Reproducible technical evidence
A local, deterministic evaluation suite measures CasePilot's structured analysis, refusal behavior, privacy masking, authorization gates, bilingual consistency, and verified-fact grounding — with no API key or network request.
73
Contract checks
73
Passed
14
Analysis fixtures
0
Network calls
Evaluation suite
100%
This is a reproducible contract benchmark, not a claim of live-model or real-world accuracy.
English and Traditional Chinese intake samples, ambiguous requests, six sensitive-request categories, and four PII patterns define expected behavior without customer data.
Exact schema, classification, replay, state-transition, privacy-boundary, bilingual, grounding, and publication checks produce the score shown above.
Owner-approved work material can contribute generalized facts to new synthetic fixtures. Raw customer media and identifiers stay outside this suite.
The public judge experience runs DeterministicDemoProvider. These results do not claim live GPT-5.6 runtime accuracy, production conversion lift, real-customer validation, or permission to publish generated content.