Methodology
What the tool actually does, the research that informs it, and the limits of what its numbers can mean.
What is being measured
IsThisOriginal evaluates the content of a written draft — its claims, mechanisms, syntheses and implications — against two baselines: what ordinary public searches retrieve, and what ordinary prompts to a base language model generate. It is not an AI-writing detector, not a plagiarism checker, not a grammar checker, and not a patent-novelty tool. It makes no claims about who or what wrote the words.
Originality here is contextual, not absolute. Every draft is compared against a corpus assembled for that analysis — topic-relevant search results and standardised model generations — never against an undifferentiated global corpus. An argument can be common in one discourse and rare in another; the report always states which corpus it measured against.
Research this system draws on
The implementation is independent: no code, datasets, prompts, tables, figures or substantial text were taken from any paper. What follows are the principles used and where they come from.
Argument rarity
The Originality Evidence Engine separates structural rarity (how the argument is assembled), claim rarity, and evidence rarity, and keeps all of them apart from logical quality — an argument can be rare and bad, or common and sound.
IsThisOriginal's argument-rarity analysis was informed by Inoshita, Omura, Yamanaka, Maeda, and Tsuji, “Argument Rarity-based Originality Assessment for AI-Assisted Writing,” arXiv:2602.01560 (2026). IsThisOriginal is an independent implementation and is not affiliated with the authors.
Originality is not effectiveness
Creativity research standardly defines creative work as requiring both originality and effectiveness, as distinct properties (Runco & Jaeger, “The Standard Definition of Creativity,” Creativity Research Journal, 2012). The tool follows this: Idea Originality and Idea Quality are computed independently and neither feeds the other. Rarity is never treated as truth or usefulness.
Semantic distance
Divergent-thinking research estimates originality by how far a response sits from its conceptual neighbours in semantic space (e.g. Beaty & Johnson, “Automating creativity assessment with SemDis,” Behavior Research Methods, 2021). Following that principle, no single cosine-similarity threshold decides anything here. Each claim is placed by its rarity percentile within the analysis's own corpus, computed under two independent methods — embedding cosine and lexical overlap — with their agreement reported. A rarity signal only one method sees is treated as weaker evidence and pulled toward the middle rather than scored as rare.
Nearest-neighbour distance, mean distance to closest neighbours and local semantic density are also computed and shown alongside each claim. These are diagnostics: they let you sanity-check a percentile against something tangible, and they do not themselves enter the score.
Combinatorial novelty
Research on atypical combinations finds that highly influential work most often pairs conventional foundations with unusual combinations of prior material (Uzzi, Mukherjee, Stringer & Jones, “Atypical Combinations and Scientific Impact,” Science, 2013). The tool therefore looks past the topic: a draft on a conventional subject can still carry an original causal mechanism, comparison, conceptual bridge, synthesis or implication — and distributed precedent (components found separately, combination found nowhere) is detected explicitly.
The two-frontier baseline
The distinctive question the tool asks: is this argument merely uncommon in wording, or actually difficult to retrieve and difficult to generate?
Public-search frontier. 10–20 queries cover the thesis, each important supporting claim, the mechanism, the concept combinations, predictions and implications, plus paraphrased and opposing framings. Each retrieved source is classified by a labelled relation — never by similarity alone — and relations carry different force, strongest first:
| Relation | Effect on rarity |
|---|---|
| Direct precedent | Decisive |
| Partial precedent | Strong |
| Same conclusion, different mechanism | Strong |
| Same mechanism, different conclusion | Strong |
| Distributed precedent (components found separately) | Moderate |
| Related but meaningfully different | Slight |
| Contradictory argument | Slight |
| No meaningful match | None |
Two rules matter more than the ordering. Containment, not agreement: a source that states your argument in order to attack it still proves the argument was already in circulation, so it counts as precedent — whether it agrees is recorded separately and never changes the score. And correlated copies count once: syndicated reprints, mirrors and rewrites of one underlying piece are collapsed into a single source before anything is weighed, so a story that travelled widely does not register as many independent precedents. Your own published page, if the search finds it, is excluded outright — a draft cannot be its own precedent.
Base-model frontier.Dozens of samples from the ordinary prompts writers actually use (“write an article about…”, “write a contrarian take”, “give an original perspective”, “suggest a Substack essay”), across escalating effort tiers, optionally spread over several models. Responses are split into atomic claims and clustered; the draft's facets are placed on the accessibility ladder, where a lower level means the idea was cheaper to generate:
Model sampling estimates accessibility, not knowledge
Scoring
Every published number is produced by deterministic, inspectable code over stored evidence. Language models supply observations — relation labels, ordinal rubric ratings, counts — and never author a score. The same evidence always yields the same numbers, and each report shows the terms behind its own totals.
Idea Originality is a weighted blend across the six facets of an argument. The weights are not equal — where a draft departs from its baselines matters as much as whether it departs:
Each facet's rarity blends three independent evidence channels — public search, the model baseline, and semantic distance within the analysis corpus — with search evidence carrying the most weight and semantic distance the least. Idea Quality is computed separately from coherence, evidential support, reasoning depth, qualification and testability. Stylistic distinctiveness is reported on its own and feeds neither score: polished writing cannot raise intellectual originality, and clichéd writing cannot lower it.
Missing evidence is never scored as a middling result
A facet the draft does not contain cannot be original
Confidence and precision
Every report carries a confidence level with its reasons stated — how much of the planned search and sampling completed, how many genuinely independent sources were found, how large the comparison corpus was, whether the two rarity methods agreed, and whether any facet went unscored. Alongside the headline number, reports show a range reflecting how far the estimate could reasonably move given the disagreement between channels, and flag when the evidence is too thin for the precision a single number implies. A low-confidence 90 is not a better result than a low-confidence 89, and the report says so rather than letting the digits imply otherwise.
The candidate point of originality must clear a genuine rarity threshold on the evidence. A draft that is thoroughly derivative is reported as having none, rather than having its least-derivative sentence promoted to fill the slot.
The heatmap vocabulary
What the scores cannot mean
- Search coverage is incomplete. Books, paywalled essays, podcasts, videos, private newsletters and non-English writing are outside every run. Absence of a match is weak evidence.
- Semantic distance is not proof of originality. Two texts can be near in embedding space and argue different things, or far apart and argue the same thing. That is why relations are judged separately and two comparison methods must agree before rarity is scored high.
- Rarity is not truth or usefulness. An uncommon claim may be uncommon because it is wrong. Idea Quality exists to keep that question separate, not to answer it definitively.
- Scores are heuristic. They should be read together with the evidence beneath them — the originality map, the precedents, the frontier analysis — never on their own.
No validation study of this system has been conducted. The scores are research-informed heuristics, not scientifically validated measurements.
Why the exact weights are not published
This page describes what is compared, which evidence counts for more than what else, and every guarantee and limitation the scoring carries — enough to understand a report, check it against its own evidence, and disagree with it specifically. The individual coefficients are not published, because a full set of weights is also a recipe for writing text that scores well, and a draft tuned to the scoring would tell its author nothing about whether the idea is actually new. Each report still shows the evidence every one of its numbers rests on: the sources, the relations, the baseline samples and the per-claim comparisons.