Browser practice uses a local solver, not JEV, and never ranks.
Y
CODEBREAKER / LEG 1
You
0 / 10
GUESSCODEEXACT / NEAR
J
CODEBREAKER / LEG 2
JEV
0 / 10
GUESSCODEEXACT / NEAR
YOUR MOVE
Make a deduction
A–F to choose · arrows to move · Enter to submit
The opponent receives only its own guesses and feedback—not your secret.
STRUCTURED EVIDENCE
Inside the decision
JEV’s candidate counts, selected features, and response timing will appear here during leg two.
Rules, fairness & keyboard controls
Choose a four-symbol code using A–F. Symbols may repeat. Both codes are locked before play. First, you break the opponent’s code; then the opponent breaks yours. Each codebreaker gets ten guesses.
Exact means the right symbol in the right position. Near means the right symbol in a different position. A symbol cannot be counted twice. Feedback does not identify which positions matched.
Fewer guesses wins. An unsolved code has comparison cost 11; equal costs draw. Resigning ends the match as a loss. Ranked human legs expire after 24 hours. The opponent’s leg is recoverable and does not time out against you.
Difficulty controls the available candidate set and deterministic features supplied to JEV. No performance ordering is assumed until measured. Local and browser-practice opponents are explicitly not JEV. A remote failure continues with a disclosed fallback and excludes that match from rankings.
Letters A–F fill the selected slot; arrow keys move between slots; Backspace clears a slot; Enter submits a complete guess. Full deduction analytics and counterfactual best guesses unlock only after the match ends. Server-held secrets protect ranked feedback; external human assistance cannot be ruled out.
EVERY GUESS, ACCOUNTED FOR
Match analytics.
Exact partitions. Counterfactual move quality. No invented thinking.
Finish a match to unlock exhaustive analysis of every turn against all 1,296 legal guesses.
Uncertainty remaining
BITS · LOWER IS LESS UNCERTAINTY
● You● Opponent · Uniform reference over consistent codes, not a behavioral prior.
Decision quality
Regret compares the selected guess with the best one-step metric over all legal guesses. Different metrics may favor different guesses.
Move-by-move deduction
Feedback partitions & counterfactual alternatives
Opponent telemetry
Full structured decision audit
Choice confidence measures the provider’s answer distribution. It is not the probability of solving or winning. Missing usage is marked incomplete, never silently treated as a complete zero-cost record.
Definitions & coverage
Verify a replay
Imported files are unranked. Verification checks internal rules, outcomes, and the original salted commitments—not server provenance.
PROGRESS WITHOUT SHORTCUTS
My record.
Results stay separated by difficulty and opponent configuration.
Performance by configuration
Ranked win streaks
Match history
VERIFIED RESULTS ONLY
The leaderboard.
Match points percentage. Ten games to establish a rank.
Community distributions
Ties use lower average penalized guesses, then more games. Local play, imported replays, and fallback matches do not rank. Community access requires a fresh verified Discord launch.
MEASURE THE OPPONENT
Benchmark lab.
The same rules and candidate engine, independently measured.
Packaged benchmark results
Loading measured results…
Reproduce the experiment
npm test
npm run test:exhaustive
npm run bench:exhaustive
npm run bench:smoke
# Explicitly authorized live-provider benchmark (requires a key):
node bench/run.mjs --live --confirm-spend --secrets=16 --max-calls=100
Deterministic selectors are baselines and ablations—not JEV. Live runs require your TypeSafe credentials and an explicit call budget. Compare JEV against the identical shortlist with deterministic selection to isolate its contribution.
See the ZIP’s reports for raw per-secret, per-turn, code-pattern, paired, and summary exports. Provider latency measures wall-clock inference; local compute timings are reported separately.