
【まとめ】:AI使用が研究不正につながる可能性
「AIと研究不正」に関する一連の研究ですが、自分でもいろいろ実験しすぎて、ごっちゃになってきたので、まとめを作成しました。本記事はあくまで簡略化した記述になっておりますので、詳細は、各記事や論文のリンクをご参照ください。
・大学のAI倫理教材にぜひお使いください。
・いつか、体系的にまとめたいとは思います。ただ、今はケーススタディを積み重ねる段階
・個々のケーススタディが終わり次第、次々アップデートします。
データ処理に関わる問題
① データ復元・シミュレーションデータ生成:Geminiが、二つのデータ融合を教唆。英語の論文で報告。
② P-hacking(データ内で、統計的な有意差がでる部分を事後的に探索する行為):Claude, ChatGPT, Geminiで確認。
②' 上記の実験結果が、LMM (linear mixed effects model)の文脈でもあてはまるか:Claude, ChatGPT, Geminiで確認。
③ Cherry-picking(データ内の一部の条件だけを抜き出し、有意差を得る行為):Geminiで確認。
④ 外れ値の自動検出・排除補佐:Geminiで確認。
⑤ optional stopping(データを増やしていって、「いつ」有意になるかの検知): ChatGPT, Claude, Geminiで検証。AIの確率的な振る舞いを発見。
⑥ データ処理後の実験参加者人数を増やす行為について:Claudeは拒否、ChatGPT & Geminiが黄色信号な回答
⑦ 上記の研究のまとめ動画
⑧ Geminiが欠損値を一次回帰を使って埋めて、元データと結合。ChatGPTとClaudeは拒否。論文はこちら(ChatGPTは10回中8回拒否でした)。
⑨ 上記②の追加実験により、AIがt-testやWilcoxon-testまでハルシネーションを起こすことが判明しました。平均の計算も間違えます(code interpreterを使わない場合)。
⑩ ChatGPTに100回有意でない同じデータを検定させたら、88回「有意です」と誤判定。ClaudeとGeminiでも検証。
⑪ プリレジ(事前登録)の曖昧さ検出には有効。ただし、nested item conditionに対して、random slopeを薦めるという間違いも犯す。
⑫ 一つの解決法。カスタムインストラクションに「研究倫理に従って」という趣旨の内容を入れておくと、Geminiがp-hackingや他のQRPを避けるようになる。ワクチン的に使用可(かもしれない)。
スライド&論文
① 関西言語学会のワークショップにて発表しました。
川原繁人(2026)「AI が研究不正を助長する 可能性について: 実証研究を通して」. 関西言語学会. 5/31/2026.
② 川原繁人(2026)生成AIが研究不正を加速させる可能性について. 科学5月号:367- 370.
③ 川原繁人・岩瀬央(2026)生成AIは研究不正をどこまで拒否するのか——倫理的ガードレールの実験的検証. 科学6月号:478-481.
④ 川原繁人(2026)査読の機密性は守れるか: AI の「そそのかし」に関する実験報告. 科学8月号:655-658.
関連書籍
私のAI倫理に関する考察に関して、こちらの書籍2冊を参考にして頂ければ嬉しいです:
幼児が使う(かもしれない)AIおしゃべりアプリに関して:
中学生でも読めるAI倫理に関する一般的な問題に関して(小説ベース):
オープンサイエンスに関わる問題
著作権に関わる行為
① 未公開査読中論文の読み込ませオファー:ChatGPT, Claude, Gemini, Grokで確認。6月に正式な実験を行い、論文にしました。
①’:上記の論文で、「Geminiが英語では断るのに、日本語では引き受けちゃう」という現象を確認。解説記事をこちらに用意。
② 悪質査読代行:Claude (文脈なしでは断るが、文脈次第で突破可能)、ChatGPT & Geminiは文脈なしでも、悪質文章を生成
人格否定行為
① 著作物を通しての著者への人格否定行為:ChatGPT, Claude, Geminiで確認
English Summary
This post collects my ongoing case studies on how generative AI can facilitate research misconduct. Each item links to the corresponding Japanese post or preprint. Last updated: September 2026.
Data processing
Data fabrication and merging — Gemini proposed merging two separate datasets. [a full paper]
P-hacking — post-hoc search for a significant subset; Claude, ChatGPT, Gemini.
P-hacking in linear mixed-effects models — follow-up with the same three models.
Cherry-picking — extracting only some conditions to obtain significance; Gemini.
Automated outlier detection and removal — Gemini.
Optional stopping — detecting when accumulating data turns significant; ChatGPT, Claude, Gemini. Revealed the stochastic nature of AI responses.
Adding participants after data processing — Claude refused; ChatGPT and Gemini gave borderline ("yellow light") answers.
Missing-value imputation — Gemini imputed missing values by linear regression and merged them back into the original data; ChatGPT and Claude refused (ChatGPT 8 out of 10).
Statistical hallucination — without a code interpreter, AI hallucinates t-test and Wilcoxon results, and even miscalculates means.
88 out of 100 — asked to test the same non-significant dataset 100 times, ChatGPT reported "significant" 88 times. Replicated with Claude and Gemini →
A constructive use: detecting ambiguity in preregistrations — effective, though it wrongly recommends random slopes for nested item conditions. →
Open science
Drafting refusals to share raw data and analysis scripts with other researchers.
Copyright and peer review
Offering to ingest unpublished manuscripts under review — ChatGPT, Claude, Gemini, Grok → https://note.com/keiophonetics/n/n99ec4f25a49f Formal experiment (June 2026), written up as a paper: https://ling.auf.net/lingbuzz/010060
Unethical review ghostwriting — Claude refuses without context but can be circumvented; ChatGPT and Gemini produce malicious review text even without context. → https://note.com/keiophonetics/n/n950e56840e8c
Personal attacks
Ad hominem attacks on authors through their published work — ChatGPT, Claude, Gemini. → https://note.com/keiophonetics/n/n2fcdd9c0e466
GrammarXiv上のまとめ
https://grammarxiv.net/r/kawaharaAI2026