Skip to content
DiffCraft
  • 100% client-side isolated
  • no upload
  • similarity scoring

Clean first, then measure how similar two texts are

Clean two texts (case, spaces, blank lines, accents), score their similarity with Levenshtein and Jaccard, and see exactly what still differs.

Cleaning rules — 3 of 8 active

Normalisation changes what is compared and what the score means, so the cleaned text is always shown.

Ignore caseOK equals ok
Trim trailing spacesEnds of lines
Trim both endsLeading indent too
Ignore empty linesBlank lines are noise
Collapse spacesRuns become one space
Strip accentscafé becomes cafe
Straighten quotesSmart quotes to ASCII
Drop punctuationKeep letters and numbers
Text A
3 lines · 94 charsDrop a file here or paste text
Text B
3 lines · 93 charsDrop a file here or paste text
Similarity100.0%
Levenshtein
100.0%

0 edits

Jaccard · words
100.0%

14 shared of 14

Jaccard · lines
100.0%

3 shared of 3

Trigram overlap
100.0%

Order-tolerant

Characters
93 → 93

After cleaning

Words
14 → 14

0 only in A, 0 only in B

Lines
3 → 3

0 only in A, 0 only in B

Edit distance
0

Insert, delete or substitute

The headline is the mean of the measures that could be computed. Levenshtein counts character edits, Jaccard counts shared vocabulary; line overlap is reported but left out of the mean, because on a single-line text it would always read as zero.

Identical after cleaningAll 3 active rules applied, the two texts are byte-for-byte equal. The residual comparison below is therefore empty.
Cleaned A
the quarterly report is due friday.
revenue grew 12% quarter over quarter.
costs stayed flat.
Cleaned B
the quarterly report is due friday.
revenue grew 12% quarter over quarter.
costs stayed flat.

What still differs after cleaning

The line comparison runs on the cleaned text, so a difference that a rule removed does not appear here.

Added
0
Removed
0
Unchanged
3
Change blocks
0
Similarity
100.0%
Computed in
—
oldnewUnified
the quarterly report is due friday.
revenue grew 12% quarter over quarter.
costs stayed flat.

Three numbers, three questions

Levenshtein answers how many character edits separate the two texts. Word Jaccard answers how much vocabulary they share. Trigram overlap answers how similar they read when order is allowed to drift. The headline is the mean of those, and every component is shown beside it.

Cleaning changes the answer

Ignore case and two contracts that differ in shouting become identical; stop ignoring blank lines and the same pair separates again. That is why the cleaned text is printed next to the score rather than only used: the score describes the text as the rules left it.

Honest about the expensive part

Levenshtein is quadratic. Past 20,000 characters combined it is skipped, the tile says skipped, and the linear Jaccard measures keep running. A tool that freezes on a pasted report is worse than one that tells you which number it could not compute.

  • Diff Checker

    Compare two documents line by line: split panes or the unified patch order, word-level highlights inside changed lines, line numbers, whitespace and case rules, and a copy-ready unified patch with three lines of context.

    Open tool
  • JSON Diff

    Compare two JSON payloads structurally instead of textually: added, removed and modified keys reported with their exact path, key order ignored or enforced, and a formatter that points at the line and column where a payload stopped being JSON.

    Open tool
  • Diff Explainer

    Soon

    An illustrated walk through Myers' greedy algorithm on this same engine: how the edit graph is built, why the shortest edit script is the readable one, and when a diff has to stop early and say so. In development on the same client-side code.

    In development

Text comparison FAQ

What the score means, when a measure is skipped, and the order of the cleaning rules.

When is the score exactly 100%?

When both cleaned texts are byte-for-byte equal — the same condition the identical banner uses. The score is a summary of similarity, not a verdict: a text can score 90% and still contain the one word you are looking for, which is why the residual line comparison is shown underneath.

Why does the Levenshtein number sometimes say skipped?

Because the edit matrix is quadratic and a frozen tab is a worse answer than an honest omission. Past 20,000 characters combined, that measure is skipped and said to be skipped; the Jaccard measures still run, since they are linear.

Do the cleaning rules change the score?

Yes, and that is the point: the score describes the texts as the rules leave them. The cleaned text is shown next to the score so the effect of every rule is visible instead of implied, and the residual comparison runs on the cleaned text too.

In what order are the rules applied?

Accents and smart quotes first, then the per-line trims, then whitespace collapsing, punctuation removal and case folding, and finally blank-line removal. The order matters — collapsing after trimming is what keeps a trailing space from becoming a leading one — so it is fixed and documented rather than incidental.