721.7B+ configurations evaluated across recorded experiments

Methodology

This page explains how we test theories, score results, and determine whether a cipher approach has been eliminated.

The K4 Ciphertext

All 26 letters of the alphabet appear at least once. Sanborn has released the plaintext at two spans, 24 characters in all. That those letters sit directly under the matching carved letters is the best-supported reading, but it rests on relayed conversations with Sanborn, not on a statement from him on the record.

O B K R U O X O G H U L B S O L I F B B W F L R V Q Q P R N G K S
· · · · · · · · · · · · · · · · · · · · · E A S T N O R T H E A S
S O T W T Q S J Q S S E K Z Z W A T J K L U D I A W I N F B N Y P
T · · · · · · · · · · · · · · · · · · · · · · · · · · · · · B E R
V T T M Z F P K W G D K Z X T J C D I G K U H U A U E K C A R
L I N C L O C K · · · · · · · · · · · · · · · · · · · · · · ·
Those 24 characters are the foundation of the scoring below. If each carved letter is enciphered in place, they also fix the key values at those positions, which is what makes whole families of cipher and key testable at all. Positions on this page count from 0, as the code does; most published sources count from 1, which puts the two spans at 22 to 34 and 64 to 74.

Scoring: Crib Matching (0–24)

For any proposed method and key, we count how many of the 24 known plaintext positions come out right. That is the score. A single random key with a 26-letter alphabet would land fewer than one of them by luck, and 6 or more happens only about once in 4,000 tries. But a search tries millions of keys and keeps the best, so scores of 6 or more turn up by chance alone, and the bands below are set with that in mind.

Scoring bands against the random-key baseline
NOISE 0–9 INTERESTING 10–17 SIGNAL 18–23 0 6 12 18 24 period 2: a random key scores 4.4 of 24 period 3: a random key scores 5.3 of 24 period 4: a random key scores 6.0 of 24 period 5: a random key scores 6.6 of 24 period 6: a random key scores 7.3 of 24 period 7: a random key scores 8.1 of 24 period 8: a random key scores 9.0 of 24 period 9: a random key scores 9.8 of 24 period 10: a random key scores 10.7 of 24 period 11: a random key scores 11.6 of 24 period 12: a random key scores 12.5 of 24 period 13: a random key scores 13.4 of 24 period 14: a random key scores 13.4 of 24 period 15: a random key scores 15.3 of 24 period 16: a random key scores 16.3 of 24 period 17: a random key scores 17.3 of 24 period 18: a random key scores 17.3 of 24 period 19: a random key scores 15.3 of 24 period 20: a random key scores 13.4 of 24 period 21: a random key scores 13.4 of 24 period 22: a random key scores 15.3 of 24 period 23: a random key scores 17.3 of 24 period 24: a random key scores 19.2 of 24 period 25: a random key scores 21.1 of 24 period 26: a random key scores 23.0 of 24 2 7 12 17 21 24 26 key period crib score
The copper line is what the best-fitting repeating key scores on average by luck alone, on a wrong guess, and it is not flat. A repeating key forces every known position that falls on the same letter of the key to take the same key value, and a search picks whichever value satisfies the most of them, so longer keys buy free matches. At period 7 the baseline is 8.1 of 24 and a real result stands out. By period 24 it is 19.2, inside the SIGNAL band, so a "significant" score there means nothing at all.

The dip is not noise. K4’s known letters sit in two blocks, at positions 21 to 33 and 63 to 73, and how those blocks line up under a given period decides how many separate key letters the known positions fall on. At period 21 they fall on only 13 and the baseline falls back to 13.4; at period 24 they fall on 19 and it climbs to 19.2. Read the chart, not the number.

The practical rule: scores are only discriminating at period 7 or below. Above that the threshold is chasing the baseline, and at 24 and up the baseline has overtaken it. The curve is computed from the real crib positions rather than transcribed; the derivation is in ops/site_builder/build.py.

Bean Constraints

In 2021 Richard Bean published constraints on the K4 keystream. The best known comes from one observation: the ciphertext letter P appears at both position 27 and position 65, and the released cribs say both decrypt to R. The rest come from comparing other pairs of known positions in the same way.

One equality, 242 inequalities
k27 = k65 P → R P → R 0 21 33 63 73 96
Same ciphertext letter, same plaintext letter, so the key value at those two positions has to be identical. Bean also listed 21 pairs that must differ because they share a plaintext or a ciphertext letter; those hold whatever alphabets are used. Doing the arithmetic in the standard A to Z alphabet gives a larger set: 242 pairs of key positions that must DIFFER, drawn faintly below the axis, plus 101 linear relations not shown. Together they cut the admissible keystreams at those 24 positions from 26 to the power 24 down to exactly 624. That count and the larger set hold only for the standard A to Z alphabet with each carved letter enciphered in place.

The full set shown above holds whichever variant is in play, Vigenère, Beaufort or Variant Beaufort, provided the cipher shifts each letter by a key value (an additive key), each carved letter is enciphered in place, and the alphabet is the standard A to Z. A lookup table, a physical overlay or a grid-based system is outside their reach. Two audits drew these limits. In August 2026 we found the constraints do not in general carry over when a transposition layer rearranges the letters. In September 2026 we found the larger set is tied to the A to Z alphabet: with the KRYPTOS alphabet used in K1 and K2, the true key values break 9 or 10 of the “must differ” pairs, though the equality and Bean’s own 21 pairs still hold. Results found to rely on them outside those limits have been withdrawn, reopened or re-checked.

Checking the 24 known key values directly shows that no repeating key of 26 letters or fewer can produce the K4 plaintext under direct positional correspondence and additive-key assumptions, on either the standard or the KRYPTOS alphabet. That is a deterministic proof, reproducible from the codebase. Some longer keys (27 to 29 letters, or 53 and more) never put two known positions on the same letter of the key, so they cannot be ruled out this way. It says nothing about a transposition layer that reorders positions, about versions that use two different scrambled alphabets, or about a non-additive cipher.

Two-System Model

Jim Sanborn stated at the Kryptos dedication that “there are two systems of enciphering the bottom text.” Scheidt stated that he “masked the English language so it’s more of a challenge” and that solvers need to “solve the technique first and then go for the puzzle.” The “two systems” quote is about the bottom half of the panel, which holds both K3 and K4, and K3 is a pure rearrangement, so how it applies to K4 alone is an interpretation. The quote is on the public record, but like the makers’ other general remarks about the method, we treat it as context, not proof. Any specific interpretation of that quote is a hypothesis, not a fact.

Pure transposition is independently impossible: the ciphertext contains only 2 E’s, but the known cribs require 3. Therefore at least one layer must be substitution. Beyond that, the architecture remains open. Current live project surfaces include:

The April 2026 audit explicitly rejected treating any single two-system story as the project’s established model. The correct public stance is: K4 likely involves a technique beyond straightforward single-layer classical encryption, but the exact composition is unknown.

What Has Been Tested

Over 597 experiments spanning 721.7B+ individual hypothesis evaluations have been run (with overlaps across experiments). Each result below holds only within its stated scope:

73-Character Hypothesis

The carved text has 97 characters. One working model proposes that 24 characters are filler letters (97 − 24 = 73 real ciphertext characters). This is a hypothesis, not a proven fact.

The strongest current public support is not the old filler-letter palette result; that evidence was retired in April 2026. The cleaner live structural observation is that the five carved Ws create a bounded segmentation hypothesis. Even that does not prove the letters are filler. The 73-character idea remains open, but it currently lacks independent, model-neutral quantitative support.

Note: The number 24 also appears in other K4 contexts (24 crib characters, Berlin’s World Clock has 24 facets, K3 chart has 24 rows). These coincidences are not evidence; many small integers recur naturally. They are documented here only because they are frequently mentioned in community discussions.

The Filler-Letter (“Null”) Palette (Retired April 2026)

A claim that 17 suspected filler letters in K4 used only seven distinct letters (the “null palette”) was once treated as a key observation. It is retained in the repo only as a cautionary historical case.

Matched controls showed the palette was not special. Rerunning the same search on shuffled versions of the text gave results like the real text’s, and the search itself did not reproduce the seven-letter pattern: that came from how the 17 positions had been picked (only from positions already holding those letters), so that letter set is not evidence about K4.

Palette constraints remain useful as a computational technique in some search programs, but the site no longer treats any retired palette identity as a cryptographic clue.

What Remains

Filler letters + a repeating-key cipher is proven impossible for ANY choice of 24 filler positions (none of them among the 24 known letters) at every repeat length from 1 to 23, on the standard or the KRYPTOS alphabet. The algebraic argument depends only on how many filler letters fall before, between and after the two known spans, and every possible split fails. (Repeat lengths 24–26 are too long to be decided either way on a 73-letter text.) If the 73-character model is correct and the remaining letters are enciphered in place with an additive key, that key cannot repeat every 23 letters or fewer: it would be, for example, a running key from an unknown source, a bespoke procedural method, a one-time key, or a repeating key of 24 letters or more.

On 2026-04-08 we ran an adversarial internal audit of the scope of our own testing and reclassified the frontier into “testable now” (bounded, reproducible campaigns we can run), “weakly testable” (requires better detection apparatus, not more compute), and “untestable under current clues” (requires new primary-source evidence). We are aware that classical cipher space is infinite and cannot be literally exhausted; this classification describes the scope of what we have tested under our specific assumptions, not a claim about K4 as a mathematical object. Full audit and record are in the internal status audit (docs/exhaustion_audit_2026_04_08.md in the research repository). Later audits, in August and September 2026, withdrew or narrowed several earlier results; open audits are tracked in docs/methodological_audits.md.

If you have an idea we have not tested, the Submit a Theory page is the direct path. We want to be wrong about anything we have classified as ruled out.

Validation Criteria

A candidate solution is not accepted unless it passes all of the following simultaneously:

How to Read Claims on This Site

Two questions decide how much weight a claim carries: how strong it is, and how thorough the search behind it was.

StrengthLevelTierWhat it means
Proven A1 Mathematical proof or complete enumeration, conditioned on stated assumptions. Disagree with the assumptions and the proof does not apply.
Exhausted B2 Every configuration in a defined space tested, all noise. Does not extend to untested variants or multi-layer combinations.
Descriptive C3 A pattern found after the fact, or a search that sampled rather than covered. Does not show how K4 was encrypted.
Open D4 A conjecture, or a method never properly tested.

Level and tier line up exactly at A/1 and B/2. Lower down they drift: C and D describe a finding, tiers 3 and 4 describe coverage.

Claims also carry a provenance tag, which is a separate question from strength. [PUBLIC FACT] is reputable reporting or a primary source, [DERIVED FACT] a deterministic consequence of one, [INTERNAL RESULT] our own output with a repro command, and [HYPOTHESIS] a claim that ships with a test plan instead of a result.

Odds quoted on this site are not corrected for the 1,073 experiment scripts the project tracks, unless stated. Across that many tests, some individually “significant” results are expected by chance. They are documented for transparency, not offered as proof.

Reproducibility

Most eliminations include a reproduction command you can run yourself. The codebase is public on GitHub. Clone the repo, install Python 3.11+, and run experiments with PYTHONPATH=src. The public copy can lag behind the working version, most stored result files are not in it, and some experiments need large inputs (such as the Project Gutenberg texts) that you must fetch yourself.