Methodology
This page explains how we test theories, score results, and determine whether a cipher approach has been eliminated.
The K4 Ciphertext
All 26 letters of the alphabet appear at least once. Sanborn has released the plaintext at two spans, 24 characters in all. That those letters sit directly under the matching carved letters is the best-supported reading, but it rests on relayed conversations with Sanborn, not on a statement from him on the record.
Scoring: Crib Matching (0–24)
For any proposed method and key, we count how many of the 24 known plaintext positions come out right. That is the score. A single random key with a 26-letter alphabet would land fewer than one of them by luck, and 6 or more happens only about once in 4,000 tries. But a search tries millions of keys and keeps the best, so scores of 6 or more turn up by chance alone, and the bands below are set with that in mind.
The dip is not noise. K4’s known letters sit in two blocks, at positions 21 to 33 and 63 to 73, and how those blocks line up under a given period decides how many separate key letters the known positions fall on. At period 21 they fall on only 13 and the baseline falls back to 13.4; at period 24 they fall on 19 and it climbs to 19.2. Read the chart, not the number.
The practical rule: scores are only discriminating at period 7 or below.
Above that the threshold is chasing the baseline, and at 24 and up the baseline has
overtaken it. The curve is computed from the real crib positions rather than
transcribed; the derivation is in ops/site_builder/build.py.
Bean Constraints
In 2021 Richard Bean published constraints on the K4 keystream. The best known comes
from one observation: the ciphertext letter P
appears at both position 27 and position 65,
and the released cribs say both decrypt to R. The
rest come from comparing other pairs of known positions in the same way.
The full set shown above holds whichever variant is in play, Vigenère, Beaufort or Variant Beaufort, provided the cipher shifts each letter by a key value (an additive key), each carved letter is enciphered in place, and the alphabet is the standard A to Z. A lookup table, a physical overlay or a grid-based system is outside their reach. Two audits drew these limits. In August 2026 we found the constraints do not in general carry over when a transposition layer rearranges the letters. In September 2026 we found the larger set is tied to the A to Z alphabet: with the KRYPTOS alphabet used in K1 and K2, the true key values break 9 or 10 of the “must differ” pairs, though the equality and Bean’s own 21 pairs still hold. Results found to rely on them outside those limits have been withdrawn, reopened or re-checked.
Checking the 24 known key values directly shows that no repeating key of 26 letters or fewer can produce the K4 plaintext under direct positional correspondence and additive-key assumptions, on either the standard or the KRYPTOS alphabet. That is a deterministic proof, reproducible from the codebase. Some longer keys (27 to 29 letters, or 53 and more) never put two known positions on the same letter of the key, so they cannot be ruled out this way. It says nothing about a transposition layer that reorders positions, about versions that use two different scrambled alphabets, or about a non-additive cipher.
Two-System Model
Jim Sanborn stated at the Kryptos dedication that “there are two systems of enciphering the bottom text.” Scheidt stated that he “masked the English language so it’s more of a challenge” and that solvers need to “solve the technique first and then go for the puzzle.” The “two systems” quote is about the bottom half of the panel, which holds both K3 and K4, and K3 is a pure rearrangement, so how it applies to K4 alone is an interpretation. The quote is on the public record, but like the makers’ other general remarks about the method, we treat it as context, not proof. Any specific interpretation of that quote is a hypothesis, not a fact.
Pure transposition is independently impossible: the ciphertext contains only 2 E’s, but the known cribs require 3. Therefore at least one layer must be substitution. Beyond that, the architecture remains open. Current live project surfaces include:
- Layered classical models: heavily tested in bounded structured families, but not globally exhausted.
- Procedural or physical constructions: still live, but only as explicit, finite, testable procedures.
- W-delimiter segmentation: deleting the five carved Ws destroys the old width-21 anomaly, but audits in May and September 2026 showed that says nothing about the Ws: replacing every W with one other letter leaves the anomaly intact, and deleting other letters spread through the text weakens or removes it too. By April 2026, 80+ direct single-layer tests had found no signal, so the idea is now admissible only as one layer of a multi-layer scheme; the interpretation is otherwise still open.
- Filler-letter (null) extraction: still possible in principle, but the old filler-letter palette result is retired and cannot be cited as evidence.
The April 2026 audit explicitly rejected treating any single two-system story as the project’s established model. The correct public stance is: K4 likely involves a technique beyond straightforward single-layer classical encryption, but the exact composition is unknown.
What Has Been Tested
Over 597 experiments spanning 721.7B+ individual hypothesis evaluations have been run (with overlaps across experiments). Each result below holds only within its stated scope:
- Repeating-key ciphers (Vigenère, Beaufort and Variant Beaufort, standard and KRYPTOS alphabets, keys of up to 26 letters): proven impossible under direct positional correspondence and additive-key assumptions (Level A). With one keyword-mixed alphabet used the way K1 and K2 used theirs, key lengths 1 to 22, 24 and 25 are ruled out whatever the keyword; versions with two different mixed alphabets are not fully ruled out.
- Self-keying ciphers (every variant, both alphabets, starting keys of up to 25 letters): proven impossible under direct positional correspondence and additive-key assumptions (Level A). Some starting keys of 27 letters or more can fit the known letters, so the proof does not cover longer keys, and self-keying combined with a letter-rearrangement layer is not ruled out.
- Many structured substitution + rearrangement combinations: many rearrangement families, each searched only within bounded settings, with no solution
- VIC-family ciphers (a Cold War Soviet hand cipher): extensively tested across multiple variants
- Letter-pair ciphers (Four-Square, Playfair, Two-Square): the standard forms need a 25-letter alphabet, but K4 uses all 26 letters; apparent high scores in other tests were overfitting artifacts
- Specialized ciphers: Feistel, Gronsfeld, Porta, and more. Gromark is only partly ruled out: its standard form (an unscrambled plaintext alphabet, key digits 0 to 9) cannot fit the known letters in place, but the form with scrambled alphabets on both sides is still open in our records, because earlier sweeps used fixed alphabets.
- Running key from 60,000+ public English texts (106 billion position checks, about half of them on 73-letter versions of the text with suspected filler letters removed): no candidate. A narrower April 2026 result, that a column-rearrangement layer at widths 6, 8 or 9 blocks every running key regardless of source text, has been disputed since August 2026 and is not relied on: it applied the Bean constraints across a transposition, where they do not hold.
- Two-layer compositions tested (105,692 branches): additive × transposition, transposition × repeating-key substitution, 6 stateful families: max crib score 6/24 within the registered layer families and default keyword sets. Some of the transposition × repeating-key branches ran on an 80-letter text made by removing the 17 letters of the retired filler pattern (see below).
- Three-layer non-columnar compositions tested (838,350 branches, April 2026): {additive, Vig, Beau} × {myszkowski, rail fence, route, block transposition} × {additive, Vig, Beau} — max crib score 7/24. This covers the enumerated layer registry with default parameter generators; non-registered outer families (e.g., homophonic, bifid, four-square as composition outer) are not included.
- Rearrangements of a 73-character text: about 4.5 million structured rearrangements tested with each cipher variant, but on a text built from a filler pattern that has since been retired (see below), so this result carries little weight
- Sculpture reading paths as keys: 10,777 paths tested, all noise. The search used a copy of K3’s plaintext that goes wrong after its first 205 letters, so paths reading the rest of K3 did not test the real text (found in a September 2026 audit).
- Grid-position-based keys: key derived from position on the grid, tested only on a 73-letter version of the text under one filler model, no signal (best 5 of 24 known letters)
- Morse code hypotheses: multiple Morse-based encoding schemes, all noise
73-Character Hypothesis
The carved text has 97 characters. One working model proposes that 24 characters are filler letters (97 − 24 = 73 real ciphertext characters). This is a hypothesis, not a proven fact.
The strongest current public support is not the old filler-letter palette result; that evidence
was retired in April 2026. The cleaner live structural observation is that the five
carved Ws create a bounded segmentation hypothesis. Even that does not prove
the letters are filler. The 73-character idea remains open, but it currently lacks independent,
model-neutral quantitative support.
Note: The number 24 also appears in other K4 contexts (24 crib characters, Berlin’s World Clock has 24 facets, K3 chart has 24 rows). These coincidences are not evidence; many small integers recur naturally. They are documented here only because they are frequently mentioned in community discussions.
The Filler-Letter (“Null”) Palette (Retired April 2026)
A claim that 17 suspected filler letters in K4 used only seven distinct letters (the “null palette”) was once treated as a key observation. It is retained in the repo only as a cautionary historical case.
Matched controls showed the palette was not special. Rerunning the same search on shuffled versions of the text gave results like the real text’s, and the search itself did not reproduce the seven-letter pattern: that came from how the 17 positions had been picked (only from positions already holding those letters), so that letter set is not evidence about K4.
Palette constraints remain useful as a computational technique in some search programs, but the site no longer treats any retired palette identity as a cryptographic clue.
What Remains
Filler letters + a repeating-key cipher is proven impossible for ANY choice of 24 filler positions (none of them among the 24 known letters) at every repeat length from 1 to 23, on the standard or the KRYPTOS alphabet. The algebraic argument depends only on how many filler letters fall before, between and after the two known spans, and every possible split fails. (Repeat lengths 24–26 are too long to be decided either way on a 73-letter text.) If the 73-character model is correct and the remaining letters are enciphered in place with an additive key, that key cannot repeat every 23 letters or fewer: it would be, for example, a running key from an unknown source, a bespoke procedural method, a one-time key, or a repeating key of 24 letters or more.
On 2026-04-08 we ran an adversarial internal audit of the scope of our own testing
and reclassified the frontier into “testable now” (bounded, reproducible
campaigns we can run), “weakly testable” (requires better detection
apparatus, not more compute), and “untestable under current clues”
(requires new primary-source evidence). We are aware that classical cipher space is
infinite and cannot be literally exhausted; this classification describes the scope of
what we have tested under our specific assumptions, not a claim about
K4 as a mathematical object. Full audit and record are in the
internal status audit (docs/exhaustion_audit_2026_04_08.md in the
research repository). Later audits, in August and September 2026, withdrew or narrowed
several earlier results; open audits are tracked in
docs/methodological_audits.md.
If you have an idea we have not tested, the Submit a Theory page is the direct path. We want to be wrong about anything we have classified as ruled out.
Validation Criteria
A candidate solution is not accepted unless it passes all of the following simultaneously:
- Crib score: 24/24
- Bean constraints: PASS, with the constraints worked out for the candidate’s own alphabet and letter order (the standard set applies only to an in-place cipher on the A to Z alphabet)
- Text quality: letter patterns must match normal English (measured by how common its 4-letter sequences are)
- Letter frequency: must match the statistical profile of English text
- Readability: must produce meaningful English with recognizable words (human review required)
How to Read Claims on This Site
Two questions decide how much weight a claim carries: how strong it is, and how thorough the search behind it was.
| Strength | Level | Tier | What it means |
|---|---|---|---|
| Proven | A | 1 | Mathematical proof or complete enumeration, conditioned on stated assumptions. Disagree with the assumptions and the proof does not apply. |
| Exhausted | B | 2 | Every configuration in a defined space tested, all noise. Does not extend to untested variants or multi-layer combinations. |
| Descriptive | C | 3 | A pattern found after the fact, or a search that sampled rather than covered. Does not show how K4 was encrypted. |
| Open | D | 4 | A conjecture, or a method never properly tested. |
Level and tier line up exactly at A/1 and B/2. Lower down they drift: C and D describe a finding, tiers 3 and 4 describe coverage.
Claims also carry a provenance tag, which is a separate question from strength.
[PUBLIC FACT] is reputable reporting or a primary source,
[DERIVED FACT] a deterministic consequence of one,
[INTERNAL RESULT] our own output with a repro command, and
[HYPOTHESIS] a claim that ships with a test plan instead of a result.
Odds quoted on this site are not corrected for the 1,073 experiment scripts the project tracks, unless stated. Across that many tests, some individually “significant” results are expected by chance. They are documented for transparency, not offered as proof.
Reproducibility
Most eliminations include a reproduction command
you can run yourself. The codebase is
public on GitHub. Clone the repo,
install Python 3.11+, and run experiments with PYTHONPATH=src. The public copy
can lag behind the working version, most stored result files are not in it, and some
experiments need large inputs (such as the Project Gutenberg texts) that you must fetch
yourself.