A statistics package or image-processing software you install stays put until you decide to upgrade it. Most AI tools do not work that way. They are changed on the vendor’s schedule, and the version answering your question this morning may not be the version you evaluated last spring. For researchers this creates a hidden risk: a workflow you validated against a set of behaviors is now running on something else.
What Tool Updates Cause Problems for Researchers
Updates arrive in 3 forms, each with different consequences:
| Type of change | What it looks like | Why it matters to you |
| Capability change | The tool can now do something new, such as search the web or read your uploaded PDFs. | Your old assumptions about what the tool cannot do are no longer safe. E.g., Tool A only corrected grammar previously, but now it can supply citations from web searches and it may add citations without you even realizing it. |
| Behavior change | Same request, different quality, tone, or level of detail in the answer. | A prompt that worked reliably may start producing weaker output in a non-obvious manner |
| Policy change | The vendor alters what it does with the text you type in, or how long it keeps it. | Your institution’s IRB approval of that tool may no longer hold. |
The practical conclusion is simple: you cannot track every release, so your safeguards have to be independent of the version.
How are Popular AI Tools like ChatGPT and Claude Changing?
Models and versions of popular LLMs change every few months, so it is more useful to think in terms of capabilities. 4 shifts matter for research work.
| Capability | What it does | The trade-off |
| Reasoning modes | The tool spends longer working through a problem before answering. | Better on hard analytical questions, but it can still produce convincing reasoning that is wrong. |
| Larger inputs | You can submit whole protocols or dozens of papers at once. | Useful for synthesis, and a much larger confidentiality exposure if the material is not yours to share. |
| Research agents | The tool searches, reads, and compiles across sources on its own. | Good for scoping an unfamiliar area, but it does not give you a search strategy anyone else can rerun. |
| Grounded retrieval | The tool pulls real documents and links to them. | The single most important difference between tools. Ungrounded answers are where invented citations come from. |
3 Questions to Ask About Any AI Tool
- Where did this answer come from, and can I open the source myself?
- What happens to the text I type in, and does my institution allow this tool for this kind of material?
- If this tool changed last month, would I be able to tell?
What Has Not Changed in Large Language Models
Capability has improved. Accuracy, reliability, and trustworthiness have not improved at the same pace. Hallucinated output still persists, and can damage a research paper’s chances of getting published.
Topaz et al.’s (2026) study published in The Lancet audited roughly 2.5 million biomedical papers and 97 million citations indexed in PubMed Central. Here’s what they found:
| Period | Papers with at least 1 fabricated reference |
| 2023 | 1 in 2,828 |
| 2025 | 1 in 458 |
| First 7 weeks of 2026 | 1 in 277 |
Of note, the fabricated references were not obviously defective. They were correctly formatted, addressed real scientific topics, were attributed to real researchers, and carried plausible dates. Most affected papers contained only 1 or 2 fake citations among several dozen, which suggests the errors were accidental rather than deliberate.
Performance also varies sharply between tools. A 2026 study in the Annals of the Royal College of Surgeons of England tested 9 chatbots, including ChatGPT-5, DeepSeek DeepThink, Google Gemini 2.5 Flash, and Perplexity Research, on common surgical questions. The worst performer, Grok 3, had 34% of its references fabricated.
Literature Search: Where AI Helps and Where It Does Not
If you’re using AI in your literature search, the most importance decision you make is not which product to buy, but which literature search tasks you are willing to hand over.
| Reasonable uses | Poor uses |
| Scoping an unfamiliar area before a formal search | Treating an AI summary as evidence you can cite |
| Finding terminology a subfield uses that you did not know to search for | Asking a general chatbot to generate a reference list |
| Tracing citations forward and backward from an anchor paper | Substituting AI output for a documented database search |
| Drafting screening criteria for later human review | Relying on AI to judge whether a study is methodologically sound |
| Summarizing a paper you have already retrieved and read | Using AI results in a systematic review without recording how they were obtained |
Need for tools that support reproducibility
A systematic review requires a search strategy that another team can rerun and reproduce. Most AI tools cannot give you that, and their results may differ next month even for an identical question. If you use an AI tool at any stage of a systematic review, record the tool, the version, the date, and the exact request, and treat its output as a starting point for a conventional search rather than a replacement for one.
AI tools for manuscript writing: the January 2026 ICMJE Rules
The International Committee of Medical Journal Editors updated its Recommendations in January 2026. The revision added a new Section 5 devoted to the use of AI in publishing. If you use AI for writing a paper, this is the change with the most direct effect on your next submission.
The Core Requirements
- AI tools and AI-assisted technologies cannot be listed as authors, because they cannot take responsibility for the accuracy or integrity of the work.
- Human authors retain full responsibility for the content regardless of whether AI was used.
- AI use must be disclosed at the point of submission, in the submitted work itself and not only in the cover letter.
- AI tools cannot be listed in the reference list as sources, because they are not considered authoritative sources of scientific information.
- AI-generated text and images must be checked for plagiarism, and failure to disclose AI use may be treated as scientific misconduct.
Patient Data Considerations for AI Tool Updates
Once patient information leaves your institution’s systems, you cannot retrieve it. Institutional policies here are stricter than most researchers assume, and de-identification of the data alone frequently does not satisfy your ethics board or information security office.
| Institution | Stated position |
| Penn Medicine | Sharing patient or research participant information with public AI services is not permissible under HIPAA or Penn Medicine policy, even when the data are de-identified. |
| Columbia University Irving Medical Center | As of March 2026, the AI chat services offered centrally were not approved for sensitive or HIPAA-regulated data. |
| Yale | HIPAA compliance is decided case by case by University and Health System committees. Vendor claims of HIPAA compliance do not meet that standard. |
| University of Chicago | Researchers using AI tools in human subjects studies must inform the IRB, which needs to know about any technology used to collect, process, or analyze participant data. |
Practical Rules
- Do not paste protected health information into a general-purpose AI tool, whether or not you believe it is de-identified.
- Just because a vendor claims to be compliant with HIPAA or similar legislation, it doesn’t mean that your institute will accept use of their tool or that funders, examiners, or peer reviewers will not object to it.
- Disclose AI tools to your IRB when they touch human subjects data at any stage.
- Remember that unpublished data from collaborators, and manuscripts you are peer reviewing, are also other people’s confidential material.
- Use synthetic or fully de-identified data when you are testing what a tool can do.
How to Check AI Output Regardless of AI Tool Updates
These 6 habits do not depend on which tool or version you are using, which is exactly why they work.
- Check every reference against PubMed or a resolving DOI, and confirm the source supports the specific claim.
- Decide before you start what the tool is allowed to touch: no patient data, no unpublished data from others, no manuscripts under review.
- Keep a short log of the tool, the version, the date, and what you asked. You will need it for the disclosure statement.
- Reread the journal’s or funder’s AI policy at submission time rather than at project start.
- Build a small test set of 5 questions you already know the answers to, and rerun it after any major update. Vendor scores for “accuracy” tell you very little about how AI will perform in your study.
- Assign a named person to review and verify every section AI has touched. Accountability cannot be delegated to software.
How to Stay Abreast of AI Policy Changes
- Bookmark 3 pages and check them once a quarter: the instructions for authors of the journals you publish in, your institution’s AI policy page, and your funder’s notice list.
- When a tool announces a significant update, spend 10 minutes rerunning your own test set before you trust it with real work.
- Ask your research group to agree on 1 written policy covering disclosure and patient data, rather than leaving each member to improvise.
- If you supervise trainees, review their AI habits directly. Many practices are inherited from supervisors whose own habits formed before these requirements existed.


Comment