Imagine submitting the manuscript you have spent years writing, only to be told that a piece of software considers it “96% AI-generated.”
What happens next?
A detector score may create suspicion, but it does not prove who wrote the book. It cannot see your early notes, abandoned chapters, editorial conversations, or the hundreds of decisions that shaped the manuscript. It sees only the finished manuscript and compares it with machine-generated patterns.
For publishers, that is no longer enough. The industry needs better ways to manage undisclosed AI-generated content, uncertain rights, and possible contract breaches. A publisher acquiring a book must have confidence that the author controls the rights being licensed and has been truthful about how the work was produced.
But the answer is unlikely to be a better guessing machine. Rather, it will be a better trail of evidence.
Publishing Now Has an Evidence Problem
Generative AI has made it remarkably easy to produce a complete manuscript. A user can generate an outline, draft chapters, imitate a genre, rewrite passages, and polish the result without leaving any obvious sign in the final output. This creates an uncomfortable problem for publishers. The manuscript may look original, but its appearance alone reveals very little about the process behind it.
The concern is not simply that AI was used. Authors already use AI in many legitimate ways, including brainstorming, language editing, translation, research organization, and checking consistency. Prohibiting every form of assistance would be both unrealistic and difficult to enforce.
The more important question is whether AI assisted the author’s creative work or substituted for it. That distinction matters because publishers do not merely buy a file containing words. They acquire or license rights from an author. Publishing agreements typically depend on assurances that the work is original, that the author has the authority to grant those rights, and that the manuscript does not infringe someone else’s intellectual property.
AI complicates every part of that arrangement.
In the United States, the Copyright Office maintains that AI-generated materials are not protected by copyright. Human-authored expression may still be protected when AI is used as an assistive tool or when a person creatively selects, arranges, or modifies generated material. The assessment, however, is made case by case. The mere presence of AI does not destroy copyright, just as entering a detailed prompt does not automatically establish authorship.
For a publisher, this creates a risk. A book written by a person who occasionally used AI to improve awkward sentences is very different from a book generated chapter by chapter and submitted under a human name. Yet both might arrive as polished Word documents with no visible indication of how they were made.
The finished manuscript cannot answer the question on its own.
Why an AI Score Cannot Settle the Matter
The publishing industry’s first instinct has been to inspect the text itself. If a machine can generate prose, perhaps another machine can recognize it.
This is the promise behind AI-generated-text detectors. They examine patterns in a document and calculate how closely those patterns resemble material produced by language models. The result is often presented as a percentage, which gives the impression of a scientific approach.
But a probability is not a finding of authorship.
Even detector companies acknowledge this limitation. Turnitin, for example, states that its AI-writing model may misidentify human-written, AI-generated, and AI-paraphrased text. It advises that the score should not be used as the sole basis for taking adverse action.
There are good reasons for that caution. Human writing can be highly predictable, particularly after several rounds of editing. Genre conventions can make prose look uniform. Formulaic nonfiction, technical explanations, instructional books, and writing by people working in a second language may all contain patterns that a detector associates with AI.
At the same time, AI-generated content can be edited, paraphrased, or combined with human writing until the original patterns become much harder to identify. Research into AI detection has found fundamental limits to what can be inferred from finished text alone, especially as language models become better at imitating human expression.
This may create an awkward situation. A genuine author may be accused because their prose appears too predictable, while someone who used AI extensively may escape attention because they edited the output more thoughtfully.
Today, journals have begun asking researchers to show the evidence behind their work. Book publishing now faces a related problem, but the evidence is different. A novelist has no laboratory dataset to produce. A biographer may have interview recordings and archival notes, but not reproducible experiments. A trade author may have years of informal research scattered across notebooks, emails, and drafts.
The question is therefore not simply whether a manuscript “sounds like AI.” It is whether the account of its creation is credible.
Watermarks Can Identify Some AI Text, Not All of It
One proposed solution is to mark AI-generated text at the moment it is created. Text watermarking works by subtly influencing the words or tokens selected by a language model. The resulting pattern is not noticeable to an ordinary reader, but a detection system designed for that watermark may be able to recognize it.
Google DeepMind’s SynthID, for example, adjusts token probabilities in text generated through supported Google AI products. This gives the output a hidden statistical signature without adding a visible label to every sentence.
This is more useful than trying to investigate the origin from writing style alone. If a known watermark is present, a publisher may have evidence that at least some of the text passed through a particular AI system.
But watermarking has obvious limits. It works only when the model provider chooses to implement it. An author could use an unwatermarked open model, an older system, or a platform with a different protocol. Substantial rewriting or paraphrasing may weaken some watermarks. A publisher would also need access to an appropriate verification system.
Most importantly, the absence of a watermark proves almost nothing. It does not establish that the text was written by a person. It only confirms that the verifier did not find the particular signal it was designed to detect.
The presence of a watermark also does not settle the question of authorship. A writer might use an AI-generated paragraph as raw material and transform it through substantial creative revision. Another might generate an entire chapter and make only cosmetic changes. The same technical signal could appear in both situations, but the human contribution would be very different.
A watermark can provide evidence about origin. But it cannot make the editorial or legal judgment that follows.
C2PA Is a Receipt, Not a Lie Detector
A more ambitious approach is emerging through the Coalition for Content Provenance and Authenticity, better known as C2PA.
C2PA is an open technical standard for attaching tamper-evident content credentials to digital files. These credentials can record signed claims about how a file was created, which tools were used, and what actions were performed on it. When supported applications preserve those credentials, they can provide a verifiable history of the asset.
This has obvious appeal in publishing. Imagine a manuscript carrying a record showing that it began in a recognized writing application, passed through several rounds of revision, was edited by another person, and later entered the publisher’s production system. Supported uses of generative AI could also be declared within that history.
Instead of asking software to guess where the text came from, the publisher could inspect claims recorded during its development.
While this may be a meaningful improvement, it is not definitive proof.
C2PA verifies that credentials are properly signed and that the information has not been altered without detection. But it does not guarantee that every statement within the credentials is true. Its core specification also does not automatically prove the personal identity of the writer. Identity systems can be connected to provenance workflows, but attribution to a specific individual is not inherent in every content credential.
Nor does C2PA guarantee a complete history. Credentials may be removed, unsupported applications may fail to preserve them, and parts of a manuscript’s development may occur outside the participating ecosystem. A document with no credentials should not, therefore, be treated as fraudulent. It may simply have been created using ordinary software that does not support them.
The distinction is important.
C2PA can help establish what happened to a file. It cannot, by itself, establish what happened inside the author’s mind.
For that reason, the publishing industry should resist calling content credentials a digital passport that settles authorship. A better analogy is a series of signed receipts. Each receipt contributes information about the journey, but the reader must still decide what that information means.
What Publishers Can Do Now
Universal provenance infrastructure may take years to develop. Publishers cannot wait for every word processor, AI platform, editorial system, and retailer to adopt the same standard.
Fortunately, they do not need to.
Publishers can build a fairer authorship-verification process from tools and evidence that already exist. The goal should not be to investigate every author. It should be to establish clear expectations and a proportionate response when credible concerns arise.
1. Ask About AI Use Before the Contract Is Signed
Publishers should tell authors what kinds of AI assistance are acceptable and what must be disclosed. The question should be specific enough to distinguish language support from the generation of substantial creative content.
A vague declaration asking whether “AI was used” will produce vague and inconsistent answers. A better policy might distinguish between brainstorming, research assistance, translation, language editing, structural revision, and direct generation of publishable text. Transparency becomes easier when the rules are clear.
2. Update Author Warranties Carefully
Publishing agreements should address AI-generated material, responsibility for factual accuracy, originality, third-party rights, and the author’s obligation to disclose material uses of generative systems.
The wording should not assume that any AI involvement makes a manuscript unacceptable or unprotectable. It should focus on whether the author can grant the promised rights and whether the use complies with the publisher’s policy.
Contracts should allocate responsibility. They should not pretend to solve technical uncertainty through aggressive language alone.
3. Treat Detection as a Signal, Not a Verdict
An AI score may justify a closer look, but it should never trigger an automatic rejection, accusation, cancellation, or withholding of payment.
Editors should first examine the manuscript itself. Are there fabricated citations? Sudden changes in voice? Repeated factual errors? Passages that appear inconsistent with the author’s earlier work or proposal? Is there conflicting information in the author’s declaration? Suspicion should begin a review, not end one.
4. Ask for Relevant Process Evidence
When a genuine concern remains, the publisher can ask the author to explain how the manuscript was developed and provide relevant supporting materials.
Depending on the book, these might include:
- Early drafts and outlines
- Version histories from Word, Google Docs, Scrivener, or other writing tools
- Research notes and source files
- Interview recordings or transcripts
- Correspondence with editors, co-authors, or research assistants
- Earlier versions shared with agents, reviewers, or colleagues
- A clear description of where and why AI tools were used
None of these items proves authorship on its own. Version histories can be incomplete, and digital activity can potentially be manufactured. But several pieces of evidence that align with one another can provide a credible account of the manuscript’s development.
5. Give the Author a Meaningful Right to Respond
Authorship disputes carry serious reputational and financial consequences. Publishers should therefore document how concerns are assessed, who reviews the evidence, and how an author can challenge a preliminary conclusion.
This is particularly important when the initial concern comes from automated software. A human decision should not merely repeat what the software reported. Fair process is not an inconvenience added to AI governance. It is part of AI governance.
What Authors Should Keep
Authors should not have to surrender their entire digital lives to prove that they wrote a book. Yet maintaining a reasonable creative record is becoming a sensible professional habit.
Keep early outlines, notes, and drafts. Preserve significant editorial correspondence. If the book depends on interviews, archival research, or original data, organize those materials so that important claims can be traced to their sources. Use version history when it is available, but do not assume it will tell the whole story.
Authors who use AI should also retain a simple record of what tools were used and for what purpose. There is a meaningful difference between asking an AI system to identify repetition and asking it to write an entire chapter. A clear record makes that difference easier to explain.
Trust Needs a Process, Not a Percentage
The dream of a perfect AI detector is attractive because it promises a simple answer. Upload the manuscript, wait a few seconds, and receive a number that tells us whether the author is telling the truth.
Publishing rarely works that neatly.
Authorship is not merely the physical act of typing words. It includes conception, selection, judgment, revision, arrangement, and responsibility. Generative AI can participate in some of those activities without taking over all of them. No percentage can reliably measure that creative relationship from the final text alone.
Watermarks may help identify output from participating AI systems. Content Credentials may provide tamper-evident records of a file’s history. Version histories, notes, and correspondence may show how a manuscript developed. Author declarations and contractual warranties can establish expectations and responsibility.
Each provides part of the picture.
The future of authorship verification will therefore not depend on one detector or one universal digital passport. It will depend on several forms of evidence, interpreted by people under policies that are transparent, proportionate, and open to challenge.
Publishers will probably never have a button that proves who wrote a book. What they can build is something more realistic and more valuable: a fair system for establishing trust from evidence rather than suspicion.