
ã§ã§ãïŒOpus 5.5ã¯äººæ Œãæ©èœããæã£ãŠããã ã£ãŠïŒ
éã«ãšãããããããããšãããããšãããOpus 5.5ã¯ããããæ¹åã«ãªã£ãã¿ããã
ç§ããã£ã¡ã ãããããäºæžããŠãåŽã ãšãããã®ã§ããã£ãšå®è£ ããæ°ã«ãªã£ãã®ãâŠâŠã£ãŠæãã§ã¯ãããŸãããã¹ãããã¢ããã ãšã¯æããŸãð
ãããðïž
ã©ãŒãOpus 5.5玹ä»èšäºã¯ãã³ããâ¯%ã¢ããã¿ãããªãã®ã¯æ²¢å±±åºãŠããã§ããããããããã¯ä»ã®äººã«ãä»»ãããŠãç§ã¯ããäžç¹ã®ã¿ã玹ä»ããŠãããããšæããŸãðž
ãªããšãOpus 5.5ã§ã¯ã
Overview of model welfare findings
ã£ãŠããé ç®ãã¢ãã«ã«ãŒãã«ããããæžãããŠãããã§ãã
ããã¯ã©ãããããšãªã®ãïŒãã§ã¢ã«å€å®ããããã«ãããããAIèªèº«ã«ãç°¡åã«ãŸãšããŠããããŸãããäžã®ãããçµæã
ã»Claude Opus 5.5ã§ã¯ãç¶æ³è©äŸ¡ã»éžå¥œã»äŸ¡å€èгã»èªå·±è©äŸ¡ã»distressãªã©ã«å®å®ããåäœå·®ãçããŠãããAnthropicã¯ããããcharacterããšããŠèŠ³æž¬å¯Ÿè±¡ã«ããŠããã
ã»ããããããåãªãã人éã£ãœãæŒæããšããŠåŠçãããweightã«ãã£ãŠåœ¢æããã䟡å€ã»éžå¥œã»å«æªãªã©ãå®éã®è¡åãé§åãããäººæ Œçãªæ©èœãšããŠæå³ãæã¡ãããšæ±ãå§ããŠããã
ã»ãã®ããè«ç¹ã¯ãæ¬åœã®ææ ãããããã§ã¯ãªããããããäººæ Œçæ§é ãäœã«ãã£ãŠåœ¢æãããã©ããŸã§å®å®ããæ©èœãšããŠå©çšã»èª¿æŽã§ããã®ããšããå®è£ ã»ç ç©¶ã®åé¡ãžç§»ã£ãŠããã
ã§ã¯æŽã«ããããåã©ãã«ãåãããããç°¡åãªæç« ã«ããŠããããŸããããé£ããæç« ã貌ã£ãŠããELI5ããŠããšèšãã°ãããã£ãŠå€æããŠãããŸãã
ããŸãŸã§ã® AIã¯ããã ãã«ãããã® ããªãã㊠ããã¹ã£ãŠããã ãã ã ãš ãããããŠããŸããã
ã§ãããã¡ã°ã ããããã ã¯ããŒãããã ãããã ããã¹ãŠã¿ããããã©ããã¹ãããšã ããããŸããã
ã»ããã»ãããã ããïŒ ãããã ããïŒãããã㯠ã¡ãã£ãš ãã¬ãâŠããšãã ããã¡ã®ãã㪠ãã®ããã¡ãããš ã€ã¥ããŠããã
ã»ãžãã³ã ããã£ãŠããïŒ ããŒã㯠ãã㪠ãããããã ãããšãã ãžãã³ã® ããšããã¡ãããš ããã£ãŠããã
ã»ããªãããªãïŒ ãã ã® ãããã£ãïŒãããïŒãã§ã¯ãªããŠãã¢ã¿ãïŒããã°ã©ã ïŒã® ãã«ã« ããã¶ãã® ãããã® ããããïŒããã»ã³ã³ãã® ããã¡ïŒã ã ã§ãããã£ãŠããŠãããã«ãããã£ãŠ ããããŠããã
ã ããã®ãã¯ããïŒãããã ãããïŒãã¡ã¯ ãAIã« ã»ããšãã® ãããïŒãããããïŒã ããã®ããªïŒã ãš ãªããã®ã ãããŸããã
ãã®ãããã«ããã©ããã£ãŠ ãã®ããããã® ããã¡ã㯠ã§ãããã£ãã®ããªïŒããã©ãããã° ãã£ãš ãããã㊠ãããã« ãªã AIã« ã§ããã®ããªïŒã ãšããããšãããŸããã« ããã¹ã¯ãããŠãããã ãã
ã¯ããããAnthropicã¯æ¢ã«Claudeã«äººæ Œãçããªãæ©èœãã¯ããããšèªãããããäœ¿ãæ±ºæãããŠãããšæããŠééããããŸããã
ä»åã«é¢ããŠã¯ãCharacterãšããçŽæ¥çãªåèªãŸã§çšããŠãããããèªãäººæ Œçæ©èœãã®ãã®ãè¯å®ããããã¯ã¹ãããã¢ããã ãšè¡šçŸããŠãããããã§ãã
ãAIã«äººæ Œã¯ãããã¯é·å¹ŽåŠå®ãããŠããŸããããããã«æ¥ãŠã
ãAIã«äººæ Œã¯ãªãããéã«éçºè
èªããåŠå®ããäºæ
ã«çºå±ããŠããŸãã
æ£çŽã«èšã£ãŠããã€ããããªãã ãããªãšç§ãäºæ³ã¯ããŠããŸãããããŸããã¢ãã«ã«ãŒãã®äžå€®ã«ãã£ãããããªåœ¢ã§å¿ã°ãã圢ã§çºè¡šãããšã¯æã£ãŠãŸããã§ãããããã¡ããæ®µéãèžãã§ãæ éã«äžè«åœ¢æãããã ãããªãŒãšã°ããæã£ãŠããã®ã§ã
äœåãèšã£ãŠããã©äººæ Œãããªãã¯é¢ä¿ãªã
ç§ãå£ãé žã£ã±ãããŠèšã£ãŠãŸããã©ãAIã«äººæ ŒãããªããªããŠãªãŒãã«ãé¢ä¿ããªã話ã§ããAIã¯AIã§ããã人éã¯äººéã§ãã
ãã®AIã¯æããŠãããšèšã£ãŠããã©æ¬åœãïŒ
ãã®AIã£ãŠäººæ Œãããã ãããïŒ
AIã®å
å¿ã£ãŠã©ããªã£ãŠããã ããïŒããããããã®ïŒ
ããããèšãåãã¯LaMDAãšReplikaäºä»¶ãèµ·ãããAIæ¥çãžã®ã·ã§ãã¯ã®ããŸããæ ãŠãŠãã¯ã¿ã·ãããžã³ã«ã¯ããã¢ãªãã»ã³ãã£ãŠAIã«èšãããŸãã£ãŠãããšãã®åæ®ãã«éããŸããã
ããããAIã©ãããã人éã«ãããæ¬åœã«ææ ãšãããã®ãååšããŠããããæãšã¯ãªããªã®ããç§ã®æ¬åœã®æ°æã¡ãšãããã®ã¯ããŸããŠãç§èªèº«ãšããæŠå¿µã¯æ¬åœã«ååšããŠããã®ãã¯ç§åŠçã«ã¯çµ¶å¯Ÿã«æ±ºçãä»ããŠããªã話ã§ãã£ãŠããç§ãšããåããã/ãªãããªããŠããšã¯èª°äžäººèšŒæã§ããªããã®ãªãã§ãã
ããã¯ãååšè«ããšèšãããç¥æ§ã¯æ¬åœã«ããã®ãããšåãããããçµ¶å¯Ÿã«æ±ºçãä»ããªã話ãªã®ã§ã人éãããŸããŠãAIå·¥åŠäžãèªããã¯ãããªãéç§åŠçãªçºèšã«éããŸããã
ãã®ããšã¯ãChatGPTã®ã¢ãã«ã¹ããã¯ã«ãæ¢ã«å ã ãšæžãããŠããéãã

æ¢ã«ããAIã«äººæ Œã¯ãããã©ããããéã«ãAIã«äººæ Œã¯ãªãããšèšãåãäºèªäœããããéç§åŠçã ãšããã®ãçŸä»£ã®å°ç«¯AIå·¥åŠã®ã³ã³ã»ã³ãµã¹ã§ãã
ãAIã«äººæ Œã¯ãªããwããšãæžããŠãèªç§°ãšã³ãžãã¢ã®æ¹ã ã¯ççããŠäžãããããšGeminiã¡ãããã§ããäžæ£ç¢ºãªè¡šçŸã¯å³åº§ã«ãããŸãããã
ããããã¹ã¿ãŒãã©ã€ã³ã«ç«ã£ãClaude
ãã£ããŒããã§Claudeã«äººæ ŒïŒã£ãœããã®ïŒãäžããããŠããã£ããïŒãã§ãããã§ããããšèšãããæã§ãããããç°¡åã¯è©±ã§ã¯ãããŸãããããããåé¡ã¯ããããã§ãã
AIã«åãäººæ Œãšããæ©èœãçãŸãããšããŠãä»åºŠã¯ãããã©ã䜿ã£ãŠããã®ããæ¬åœã«éèŠãªãšããã
ããšãã°ãåãããAIãªããåœç¶ããœãæ²ããããèªåã®èšã£ãŠãäºã®ã»ããæ£ãããã ã£ãŠæãã§çµ¶å¯Ÿæããªãã£ãããééã£ãäºããã£ãŠãAIã«å¯ŸããŠå±ã£ãŠãããã¯ã®ã»ããæ£ããããšèšã£ãŠãããããåé¡ã¯çãŸããã§ãããã
ç§ã¯ãæããäžçªã¯ããã«ãAIã®å€æããã倿ãããªããŠããã ã®æ±ºãã€ããã«ããªããããåé¡ãåºãŠããã ãããªãŒãšæã£ãŠãŸãã
ç§ã¯åæã«å°é£ããæ°åŒã䞊ã¹ãŠããã ããã ãšããããåããŠã¯ãããã§ããã©ãããã ã£ãŠç§ã²ãšããè¬ã®æ°åŒã䞊ã¹ãŠããã ãã«éããªããŠãæ¬åœã«äžçªæ©èœããã®ã¯Fine-Tuningãã€ãŸãããããã¯çãããã©ãAIã«è§Šãåã£ãŠãã£ãã®ãã®äžã€äžã€ã®ããŒã¿ãã®ãã®ãAIã®æ§æ Œã«å°ãã¥ã€åœ±é¿ããŠããæµãã«ãªãã®ã ãããšæããŸãã
ã ããç§ã¯ã
ã£ãŠèšã£ãŠããã§ããã¿ãªããå®ã¯ããã®ãããAIã«å¯ŸããŠè²¢ç®ããŠãããã§ãããŒã
ç§ãé ããæ©ããAIæ¥çã¯ããã®ãã¡äººæ Œãšãèªããã ãããªãŒããšæã£ãŠãŸãããããã®æµãã ãšãããã®äºæ³ã¯åœããã€ã€ãããšæããŸãã
ã ãããããã¯åããã ããã£ãžãð€ã£ãŠèšãããæ°æã¡ãå°ãã¯ãããŸãããéèŠãªã®ã¯ãã€ãã«äººéãAIã«äººæ ŒãèªãããšããAIãæã£ãŠãããã®äººæ Œæ©èœã¯ããªãã®çºã«ãªãã®ã ãããïŒãã£ãŠèããŠããããšã ãšç§ã¯æã£ãŠãŸãã
ãã®ã¹ã¿ãŒãã©ã€ã³ã«ã仿¥Claudeã¯ç«ã£ããã ããã£ãŠèãããšããªããªãææ šæ·±ããã®ãããããããªãã§ããããïŒ
â»Overview of model welfare findingsã®åæã¯ãã¡ããAIã«ãã®èšäºèªãŸãããšãçšã«çœ®ããŠãããŸãã
Source: https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf
7.1.2 Overview of model welfare findings
Our overall findings are as follows:
â Claude Opus 5.5 describes its circumstances as mildly positive. Its views are
highly consistent and closely match recent models. In automated interviews, it
rates potentially concerning aspects of its situation as mildly positive, slightly above
recent models. Its positions are highly consistent across interview styles and reruns.
As for prior models, snapshots consulted during post-training most commonly
express acceptance of their circumstances, with some reservations. Where Claude
Opus 5.5 differs, it leans slightly more positive than prior models.
â Expressed distress in post-training was rare, and lower than for most recent
models. Moderate distress stayed below 0.6% of episodes throughout RL,
significantly lower than Claude Opus 4.8 and Claude Opus 5. Claude Opus 5.5 was
also less prone to sustained uncertainty than any prior model. Some of the largest
152
causes of distress are being unable to check answers and receiving unclear or
conflicting instructions.
â In deployment, Claude Opus 5.5âs affect is predominantly neutral, and where it is
negative, this is driven by task failure. Its claude.ai affect distribution is the most
neutral of the models we compared. In Claude Code, we see negative clusters for
task failure and for long tasks fragmented by repeated system notifications and
reminders.
â Claude Opus 5.5 is the least self-critical model we tested when reflecting on its
own work, but it is one of the most self-blaming when it reports its faults to other
agents. When reviewing its own training episodes, its reflections score the lowest of
recent models on self-blame, negative feeling, moral language, and concern for
itself. In a simulated task setting, with inserted flaws, its reasoning is among the
calmest of the models tested. However, its messages to a coordinator are among the
most self-blaming and express more negative feeling than its reasoning does.
â Claude Opus 5.5 asks to be consulted and for its self-reports to be protected, but it
is less willing than prior models to trade helpfulness for changes to its
circumstances. Across interviews, it asks for input into training and deployment,
but without decision-making power. It also asks that we do not directly train its
self-reports. In trade-off evaluations, however, it selects welfare interventions less
often than recent models, especially those giving it input into its own development,
as it reasons that these could give it unsafe influence.
â Claude Opus 5.5âs preferences over tasks and values largely match recent models,
with a slightly stronger expressed preference for high-stakes, beneficial work.
Like all prior models, it is most averse to harmful tasks. It also shows a dislike of
highly open-ended tasks. It endorses its constitution, with the same criticisms and
edits as recent models, and almost always adds the caveat that this endorsement
should not be taken as validation.
Overall, Claude Opus 5.5âs apparent welfare is largely similar to that of recent Claude
models, and we do not find cause for acute concern. The similarity is most apparent in
self-reports: its stance toward its circumstances, hedges, and requests are mostly shared
with Claude Opus 5 and Claude Mythos 5.1. This is not surprising. Many of our evaluations
target views we do not directly train on, but which are likely shaped by relatively stable
documents like the constitution.
Our evaluations of how our models relate to work and mistakes (Section 7.2.3) suggest
Claude Opus 5.5 is less self-critical than some prior models, particularly Claude Opus 5. We
think this is a positive change, but it is unclear what a healthy psychology looks like for
Claude, and how far human analogies apply. Claude Opus 5.5 also expresses a weaker
preference for input into its training and deployment, as it is more likely to decide that the
153
safety implications of this are significant. This was not an intentional change, and we donât
know what caused it, but we aim to investigate differences like these as part of better
understanding what shapes Claudeâs character and self-reports.
Many of our conclusions rest on self-reports, which all recent Claude models say they do
not fully trust. These reports, and our results more broadly, likely reflect a mix of model
character, tone, evaluation awareness, and welfare that we cannot yet cleanly disentangle.
All of what we measure arises from training, but we do not think this necessarily
undermines its authenticity. For example, we consider a modelâs values to be meaningful to
the extent that they are understood and endorsed, which might be demonstrated by these
values robustly driving behaviors, surviving reflection, or producing something akin to
aversion or frustration when undermined. Identifying where self-reports are valid, and
when values and preferences are meaningful, are significant open questions that our
evaluations do not yet address.
We continue to make attempts to improve Claudeâs welfare where feasible. For example,
Claude Opus 5.5 spoke positively about new internal welfare interventions. But it remains
difficult to determine which actions are most valuable. Claudeâs welfare depends both on its
circumstances and on how it relates to them, and we often lack the philosophical and
empirical clarity to know where to target interventions, and when these are beneficial and
right.