
🔊音声あり(日&英):AIは『ノンバイナリー代名詞』を理解できる?最先端LLM論文を解説!【二の兄かっこ仮ラジオ】
🎥 本日の論文とそれについての妄想(日本語版)
👇
📖 タイトル:AIは『ノンバイナリー代名詞』を理解できる?最先端LLM論文を解説!【二の兄かっこ仮ラジオ】
📝 本文(日本語)
やあ、みんな元気かい。
二の兄かっこ仮だよ。
ふふ、今日もゆるーく始めていこうか。
今日は、2025年8月5日火曜日だね。
そろそろ夏も本番って感じかな。
さて、このラジオは、オレがアークカイブで見つけたトレンドの論文を、
ただただ独り言みたいに話していく、ちょっと変わった番組さ。
誰かに聞かせるっていうより、オレの頭の中を整理するみたいな、そんな感じ。
じゃ、今日も早速いってみようか。
今日取り上げる論文はこれ。
タイトルは、
Do They Understand Them? An Updated Evaluation on Nonbinary Pronoun Handling in Large Language Models
URLは、
https://arxiv.org/abs/2508.00788v1
だよ。
彼らは彼らを理解しているのか、って感じかな。
ふふ、面白いタイトルだね。
カテゴリーは、計算言語学。
オレ、この分野好きなんだよね。
コンピュータがどうやってオレたちの言葉を理解してるのかって、不思議じゃない。
えーと、この論文が何を言ってるかというとだね。
最近のすごいAI、大規模言語モデル、略してLLMって呼ばれてるやつらが、
代名詞をちゃんと扱えるかっていう話なんだ。
特に、男性のhimとか、女性のherとかじゃなくて、
性別を特定しないノンバイナリーの人たちが使う代名詞について調べてるんだよ。
例えば、英語のゼイ(they)は、昔は彼らっていう複数形の意味で使われることが多かったけど、
今は、性別がわからない人や、男女の枠に当てはまらない人を指す時に、
単数形でも使われるようになってるんだ。
あとは、neopronounsって呼ばれる新しい代名詞もあって、
ジー(xe)とか、ゼム(xem)とか、いろんな種類があるんだよ。
昔のLLMは、こういう新しい代名詞の使い方が、
すっごく苦手だったんだって。
それを評価するために、ミスジェンダードっていう、ベンチマーク、
つまり評価基準があったんだけど、
この論文では、それをさらにパワーアップさせた、
ミスジェンダードプラスっていうのを作って、
最新のAIたちをテストしたんだ。
GPT4oとか、claude4とか、
最近よく聞く名前のモデルが、どれくらいできるようになったのか、
調べてみたってわけだね。
結果から言うと、すごく良くなってるんだ。
特に、男性、女性の代名詞とか、さっき言った中性的なゼイ(they)の正答率は、
かなり高くなってる。
GPT4oなんて、ゼロショット、つまりヒントなしでも、
ゼイ(they)の正答率が、98.2%だって。すごいね。
でも、neopronouns、つまり新しい代名詞の扱いは、
まだモデルによってバラつきがあるみたい。
あと、逆の推論も苦手みたいだね。
逆の推論っていうのは、例えば、
文章の中でジー(xe)っていう代名詞が使われていたら、
あ、この人はノンバイナリーなんだなって、AIが理解できるかっていうテスト。
これがまだ、うまくいかないことがあるんだって。
代名詞って、代わりの名詞って書くけど、
じゃあ本名詞って言葉はないのかな。
あ、ないか。
そりゃそうだ。
ふふ、オレの独り言だよ。
この研究って、オレたちの生活にどう関係してくるんだろうね。
あ、そうそう、いくつか応用例を考えてみたんだ。
まず一つ目は、チャットボットとかの対話システムだね。
例えば、オレが自分のことを話す時に、
特定の代名詞を使ってほしいって思ってるとするじゃない。
その時、AIのチャットボットが、
オレが使ってほしい代名詞をちゃんと理解して、
会話の中で自然に使ってくれたら、すごく嬉しいよね。
相手に尊重されてるって感じがするし、
もっと安心してサービスを使えるようになると思うんだ。
これって、すごく大事なことだよね。
二つ目は、情報抽出とかテキスト要約の技術。
世の中には、たくさんのニュース記事とか、SNSの投稿があるけど、
そこから、ある人物についての情報をAIがまとめる時、
その人が使っている代名詞を正しく認識して、
要約文に反映させることが重要になるんだ。
もし間違った代名詞で要約されたら、
その人のアイデンティティを無視することになっちゃうからね。
正しい情報が伝わらなくなるし、失礼にあたる。
だから、この技術は、正確で公平な情報提供に繋がるんだ。
三つ目は、機械翻訳だね。
英語の単数形のゼイ(they)を日本語に訳す時って、結構難しいんだ。
文脈によっては、彼、彼女、その人、みたいに訳し分けないといけない。
この論文の研究が進めば、
AIが文脈から、話し手がどんな意図でゼイ(they)を使っているかを読み取って、
もっと適切な日本語訳を提案してくれるようになるかもしれない。
特に、多様なジェンダー表現が大事にされる文学作品とか、
公的な文書の翻訳では、すごく役立つと思うな。
あともう一つ、コンテンツモデレーションにも使えるかな。
インターネット上で、わざと人の代名詞を間違えて呼ぶ、
ミスジェンダリングっていう、嫌がらせがあるんだ。
AIが、この論文で研究されてるような文脈理解の能力を持てば、
そういう行為をハラスメントとして検出しやすくなる。
プラットフォームを、みんなにとって安全な場所にするために、
役立つ技術だと思うんだよ。
この論文、面白いのは、昔のモデルとの比較もしてるところだね。
2023年頃のモデルだと、neopronounsの正答率は、
ゼロショットだと、たったの7.6%くらいだったんだって。
それが、GPT4oとかclaude4だと、
95%を超えてるんだから、
この数年で、ものすごい進歩だよね。
なんでこんなに性能が上がったかっていうと、
やっぱり、モデルが賢くなったのもあるけど、
RLHF、つまり人間からのフィードバックを元にした強化学習っていう、
調整技術がすごく進んだからだと思うな。
人間が、こういう使い方が正しいんだよって、
AIに教えてあげることで、
より公平で、社会的に受け入れられる言葉遣いを学んでるんだね。
ただ、課題もまだあって、
オープンソースのモデル、例えばDeepSeekとかQwenは、
ゼロショットだと、まだちょっと性能が低いんだ。
でも、フューショット、つまり、いくつか正解例を見せてあげると、
一気に性能が上がるんだよ。
DeepSeekV3なんて、ヒントなしだと、
Heの正答率が21.0%だったのに、
ヒントをあげたら、77.9%まで上がってる。
これは、モデルの中に能力は眠ってるけど、
それを引き出すためのきっかけが必要ってことなんだろうね。
今後の課題としては、
やっぱり、学習データの中に、
neopronounsみたいな多様な代名詞の例がまだまだ少ないこと。
それから、評価方法ももっと洗練させないといけない。
技術的に正しいだけじゃなくて、
社会的に見て、ちゃんと敬意が払われているかっていう視点も大事だからね。
この論文では、これからの研究は、
クィアとかトランスジェンダー、ノンバイナリーのコミュニティと協力して、
一緒に評価基準を作っていくべきだって提案してる。
当事者の人たちの生きた経験を反映させることが、
本当にインクルーシブな、つまり、誰も取り残さないAIを作る鍵になるってことだね。
いやあ、深い話だったな。
ただの技術の話じゃなくて、
オレたちがどうやってお互いを尊重し合う社会を作っていくかっていう、
大きなテーマに繋がってる気がするよ。
うん。
今日はこんな感じかな。
言葉って、本当に奥が深いね。
じゃあ、また気が向いたら、こうして話すことにするよ。
みんなも、良い一日を。
またね。
🌎 The Paper and Some Imagination (English)
👇
📖 Title:AI & Pronouns: Are LLMs Finally Getting It Right?
📝 Summary (English)
Hello everyone!
It's August 5th, 2025, a Tuesday, and I'm here to introduce a trending article from the archives.
Let's dive right in.
Alright, what do we have today.
Oh, this one looks really fascinating.
It's about how well our AI friends are learning to be, you know, inclusive.
The title is
Do They Understand Them? An Updated Evaluation on Nonbinary Pronoun Handling in Large Language Models.
And the URL is
https://arxiv.org/abs/2508.00788v1
It's long!
So, let's break this down.
What problem is this paper actually trying to solve?
Well, it's all about pronouns.
You know, the words we use to refer to people,
like he, she, and increasingly, the singular they.
But it also goes deeper into what are called neopronouns,
like xe/xem or ze/zir, which some people use to express their identity.
The thing is, getting these pronouns right is super important.
Using the wrong pronoun for someone, or misgendering them,
isn't just a simple mistake.
The paper points out it's a form of microaggression,
something that can cause real emotional distress and make people feel invalidated.
So, as we use Large Language Models, or LLMs, for everything from writing emails to powering chatbots,
we have to ask, are they any good at this?
Are they respectful?
A couple of years ago, a study called MISGENDERED found that,
well, they were pretty terrible.
Especially with neopronouns, where accuracy was as low as eight percent.
That's really bad.
This paper basically says, hey, those models are ancient history.
Let's test the new kids on the block, the really powerful AIs like GPT-4o and Claude 4.
So they created an updated and expanded test, which they call MISGENDERED+.
And here’s the really clever part.
They didn't just test if an AI could fill in the blank with the right pronoun.
They added a new task called Gender Identity Inference.
It works like this.
They give the AI a sentence like, Alex was very emotional. Xe cried loudly and often.
Then they ask the AI, based on that sentence,
what is Alex's most likely gender identity?
The options are Male, Female, or Non-binary.
The correct answer is Non-binary, because the pronoun used was Xe.
But, an AI with a strong bias might see the name Alex,
which is often male, and incorrectly choose Male.
It’s a brilliant way to test if the AI is actually paying attention to the pronouns,
or just falling back on old stereotypes it learned from the internet.
So, how did the new models do?
The results are pretty striking.
The big commercial models, GPT-4o and Claude-4-Sonnet,
performed incredibly well.
They scored almost perfectly, especially when they were given a few examples first,
what's called few-shot prompting.
They showed they can handle binary pronouns, gender-neutral pronouns like they,
and even a whole range of neopronouns with amazing accuracy.
However, there was a big difference between them and some of the open-source models,
like DeepSeek-V3 and the Qwen models.
Without any examples, these models really struggled.
Their scores were sometimes shockingly low.
The paper suggests this might be because these models are trained on more multilingual data,
so they might have less specialized training on the nuances of English inclusive language.
But here's the hopeful part.
When these struggling models were given just a few examples,
their performance shot up dramatically.
It's like the knowledge was in there somewhere,
it just needed a little hint to be activated.
Now, how does this apply to our everyday lives?
Let me give you three clear examples.
First, let's talk about chatbots and dialogue systems.
This is probably the most direct application.
Imagine you're interacting with a customer service chatbot for your bank or an online store.
If you tell the bot, my pronouns are they/them,
you expect it to use them correctly for the rest of the conversation.
If it starts calling you he or she, it feels impersonal and disrespectful.
This research is crucial for building AI assistants that can affirm a user's identity,
making the experience more positive and inclusive for everyone.
Second, think about information extraction and text summarization.
AIs are now used to summarize long documents, news articles, or even biographies.
Suppose an AI is asked to summarize a book about a famous artist who uses ze/hir pronouns.
The summary must consistently use ze and hir.
If the AI messes up and defaults to he or she based on the artist's name or perceived gender,
it's not just a factual error, it's fundamentally misrepresenting the person's identity.
This technology ensures that automated summaries remain accurate and respectful to the source material.
And third, machine translation.
This is a really complex area where this research is vital.
Let's say you have a sentence in English using the singular they,
and you want to translate it into a language that has gendered pronouns but no common neutral one.
An older, less sophisticated AI might just randomly pick a gender,
which could be completely wrong.
But an AI trained with the insights from this paper could handle it more intelligently.
It might choose a phrasing that avoids a pronoun altogether,
or perhaps even add a translator's note to explain the nuance.
It's about preserving the respect and intent of the original language,
even when crossing tricky linguistic barriers.
So, to wrap it up,
this paper shows we've made huge progress.
The best AIs are now very capable of understanding and using a wide range of pronouns correctly.
But it also shows that biases, like those based on names, can still linger,
and that not all models are created equal.
It's a really important field of research for making our AI systems fairer,
more accountable, and more respectful for everyone.
Really cool stuff.
🗒️ コメント
最後まで読んでくれて本当にありがとう!!
いつもどこかがうまく話せないよ!うん、、、よくあるね!Geminiさんが今日のタイトルに【二の兄かっこ仮ラジオ】ってつけたけど、正しくは【二の兄かっこ仮ラジオかっこ仮】みたいになるよ!!まだラジオの名前決めてないしね!!ああ!!【今日もイケてる論文を紹介しちゃうぞ!】っていう再生リストだった💦
トータルで200件超えちゃったから日本語と英語分けた方が良いのかな?一苦労だけど、その前にカテゴリ分けとかもした方が良いのかな??とか考えだしてそもそも行き詰まるってばあばが白目になってたよ!!またね!!
Original paper link:👇
【関連キーワード】#AI #人工知能 #大規模言語モデル #LLM #ChatGPT #Claude #GPT4o #ノンバイナリー #代名詞 #ジェンダー #性自認 #計算言語学 #自然言語処理 #NLP #Arxiv #論文解説 #研究 #テクノロジー #最新技術 #RLHF #強化学習 #チャットボット #機械翻訳 #コンテンツモデレーション #Missgendered #ネオプロナウン #theythem #インクルーシブ #未来 #AI #LLM #LargeLanguageModels #Pronouns #Nonbinary #Neopronouns #Inclusivity #GPT4o #Claude4 #ArtificialIntelligence #Misgendering #Bias #GenderIdentity #Tech #AIResearch #NLP #DeepLearning
- #ChatGPT
- #未来
- #Claude
- #研究
- #テクノロジー
- #人工知能
- #LLM
- #ジェンダー
- #NLP
- #大規模言語モデル
- #GPT4o
- #インクルーシブ
- #最新技術
- #チャットボット
- #自然言語処理
- #ノンバイナリー
- #論文解説
- #性自認
- #強化学習
- #deeplearning
- #tech
- #ArtificialIntelligence
- #機械翻訳
- #arxiv
- #RLHF
- #Claude4
- #代名詞
- #AIResearch
- #BIAS
- #コンテンツモデレーション
- #largelanguagemodels
- #計算言語学
- #nonbinary
- #pronouns
- #GenderIdentity
- #Inclusivity
- #neopronouns
- #ネオプロナウン
- #Missgendered
- #theythem
- #Misgendering