Key Takeaways
- Funder rules target authorship, not tool use: No current policy bans AI grant writing. What policy bans is submitting a scientific argument you did not author and cannot defend.
- Disclosure requirements differ by funder: Wellcome requires a declaration, NSF encourages one, UKRI expects one for substantive use only, and NIH sets none at all.
- Confidentiality is the underrated risk: Funders state plainly that they cannot guarantee what happens to text entered into a third-party tool, and an unfunded proposal is unpublished work.
- Verification is the actual work: Fabricated references and claims that outrun their evidence are what sink proposals, and neither is visible on a read-through.
Getting caught is the least of the risks when you use AI to draft a grant proposal.
Funders have spent three years writing rules for this, and most of them aren't about detection. They are about who authored the scientific argument, whether the claims inside it are real, and what happened to your unpublished ideas after you pasted them into a text box.
An analysis of more than 125,000 NIH and NSF proposals submitted between 2021 and 2025, including confidential unfunded submissions, found that proposals showing heavy language model involvement were less semantically distinctive from recently funded work.
At NIH, those proposals were roughly four percentage points more likely to be funded and produced modestly more publications, with no advantage in citation impact. The tools helped applicants sound like the field, without any measurable gain in the impact of what followed.
Researchers getting durable value from these tools have narrowed what they ask of them. AI handles language, structure, and compliance. The scientific argument stays with the applicant, and every claim gets checked against a source before it leaves the building.
Successfully (and ethically) using AI for drafting proposals starts with knowing where funders draw the line, what their policies actually require, and how to review an AI-assisted draft before submission.
Where AI Genuinely Helps in Proposal Drafting
The clearest guide to what is safe in grant proposal drafting comes from UKRI, which sorts AI use into two categories and tells applicants outright that one of them does not need to be declared at all, which includes translating an application into English, improving the standard of English, formatting, and reducing word count.
Language, Structure, and Length
These are the tasks where the output is checkable at a glance and where nothing about the science changes.
A model can tighten a long rationale, flatten inconsistent tense, or cut a section to a page limit without changing what the section claims. For researchers writing in a second language, this is the difference between a proposal that reads as careless and one that reads as competent, and it is explicitly sanctioned.
Compliance and Formatting Checks
Reading a draft against a call document is pattern matching, which is what these systems do well.
Required sections, page and character limits, whether every aim has a matching milestone, or whether the budget justification covers every line in the budget. Treat the output as a list of things to look at rather than a list of things that are fixed. A missed requirement is still your problem at submission.
Pressure-Testing an Argument You Already Wrote
The highest-value use is also the furthest from generation.
Give a model your finished aims and ask what a hostile reviewer would attack, including:
- Which claim rests on a single preliminary figure
- Which alternative explanation you have not addressed
- Which aim fails if aim one fails
In this scenario, you are not asking it to fix anything. You are asking it to find the questions you have stopped being able to see.
There is one cost worth naming. That same analysis of federal proposals found that the more heavily a model was involved, the more closely a proposal resembled work the agency had already funded. Language models are trained to produce the expected next phrase, and a funding proposal is supposed to argue for something the field has not done yet. The more drafting you delegate, the harder that tension pulls against you.
researchers
Where the Ethical Line Sits
Most arguments about ethical AI use in academic writing stall on the question of how much AI is too much.
Funders have answered a different question. NIH states that applications, or sections of applications, substantially developed by AI will not be considered the original ideas of applicants, and warns that AI use can produce plagiarism, fabricated citations, and other forms of research misconduct.
UKRI states that applicants must not generate an entire application, or sections of one, without human involvement.
NIH has not defined what "substantially" means, and no percentage or word count exists to hide behind. That ambiguity is uncomfortable, and it also points to the workable test. Could you defend every claim in this proposal, unprompted, in a conversation with a program officer? Any passage where the answer is no is a problem, no matter how it was produced.
Four things stay with the researcher on any reading of current policy:
A second question tends to fill the space policy leaves open: will anyone know? Detection is a poor foundation for that decision, either way.
In 2023, Stanford researchers found that widely used GPT detectors consistently misclassified writing by non-native English speakers as AI-generated while correctly classifying native writing, and that simple prompting bypassed them.
NIH says it will continue to use detection technology on applications. Neither fact tells you anything reliable about your own submission. The only position that holds up is that you can defend your authorship, regardless of how a detector errs.
What Funders and Institutions Currently Expect
Policy has converged on two points and split on a third.
Every major funder holds the applicant responsible for the accuracy and authenticity of the submission. Every major funder bars reviewers from putting confidential applications into AI tools. AI disclosure is where policy diverges.
Wellcome requires a declaration on the application form, with exceptions only for translation and improving the standard of English. NSF encourages proposers to indicate the extent to which generative AI was used in the project description. NIH’s notice contains no disclosure requirement at all, as its rule governs the originality of the ideas.
The table below sets out what each funder has written down, for applicants and for the people assessing them.
The reviewer column is the more uniform half, and it has been getting firmer. The ERC’s March 2026 guidelines rest on two principles, non-delegation and confidentiality, which is the same position NIH, NSF, UKRI, and Wellcome reached separately.
Where declaration is asked of applicants, uptake is partial. Funders reported declared AI use of 25% among applicants to Wellcome’s Early Career Awards, 16% at the British Heart Foundation, and 12% at Cancer Research UK. Wellcome has said it does not pass those declarations to grant committees, so reviewers cannot speculate about how much AI a bid involved.
The takeaway? Check your institution as well as your funder. Research offices increasingly ask for a certification at routing that the application is the applicant’s own original work.
How to Verify AI-Assisted Claims Before You Submit
Two failure modes matter, and only one of them announces itself.
In one controlled test, GPT-4o produced 176 citations across six mental health literature reviews, and close to 20% referred to no real publication. Fabrication rates ran highest on the more specialized topics, which is the direction that matters for a grant proposal.
The outright fabrications are findable if you go looking. The second category is worse, because the paper exists, the authors are right, the year is close, and the identifier resolves to something adjacent.
Fabricated references also reach print. An audit of 2.5 million papers in PubMed Central’s Open Access subset identified 4,046 references pointing to publications that do not exist, with the rate of affected papers rising roughly twelvefold since 2023. The authors attribute those fabrications to a mix of paper mills, misconduct, and uncritical use of AI writing tools. Reviewers know this now, and a single unverifiable reference invites them to doubt everything around it.
A successful approach should include three passes using this specific order:
- Confirm every reference exists: Look each one up in PubMed, Crossref, or the publisher, not in the tool that produced it. Asking a model whether its own citation is real produces another confident answer, not a check.
- Confirm each source says what you say it says: Open the paper. A real citation attached to a claim it does not support is the version of this problem that survives a reference check and reaches a reviewer intact.
- Confirm your claim does not outrun your evidence: Take each assertion in the aims and ask what specifically supports it, in what system, under what conditions. Smoothed prose has a way of turning "consistent with" into "demonstrates."
That third pass is the one no reference checker performs and the one a reviewer performs first. It is also the pass that scales worst, because it means reading every cited paper against your own argument.
This is the layer QED Science works on, breaking a paper or grant into its individual claims and showing where the evidence holds up and where it does not, testing each claim against the literature rather than only against the sources you cited, without training on what you upload. The same claim-level approach underpins a quality metric that scores manuscripts on originality and validity after author names and affiliations are stripped out.
Build a Proposal That Holds Up Under Scrutiny
Funder policy will keep moving, and the six-month-old version of any of these rules is already the wrong one to work from.
The rising use of AI technology is why funders keep restating a rule that was already in place: accountability sits with the named applicants, and it is a property of the argument, not of the toolchain that produced the prose around it.
That makes the practical standard stable enough to use today. Put AI on the tasks where you can check the output at a glance. Keep the science yours and test every claim against a source before someone else does.
QED Science applies the same standard at scale, most recently in the largest blind quality assessment of preprint science conducted to date.
Upload a grant or a manuscript and see which of its claims hold.
FAQs
researchers
researchers

