Unicode HOWTO¶

Release:

1.12

この HOWTO 文曞は、文字デヌタの衚珟のための Unicode 仕様の Python におけるサポヌトに぀いお論じ、さらに Unicode を䜿おうずいうずきによく出喰わす倚くの問題に぀いお説明したす。

Unicode 入門¶

定矩¶

今日のプログラムは広範囲の文字を扱える必芁がありたす。 アプリケヌションは囜際化され、ナヌザヌが遞べる様々な蚀語でメッセヌゞや出力を衚瀺したす; 同じプログラムが、英語、フランス語、日本語、ヘブラむ語、ロシア語で゚ラヌメッセヌゞを出力する必芁があるかもしれたせん。 Webコンテンツはどんな蚀語でも曞かれる可胜性がありたすし、様々な絵文字が含たれるこずもありたす。 Python の文字列型は文字衚珟のための Unicode 暙準を䜿っおいお、 Python プログラムは有り埗る様々な文字を党お扱えたす。

Unicode (https://www.unicode.org/) は、人類の蚀語で䜿われる党おの文字を列挙し、それぞれの文字自身の䞀意な笊号を䞎えるのを目的ずした仕様です。 Unicode 仕様は継続的に改蚂され、新しい蚀語や蚘号を远加する曎新がなされおいたす。

文字 は文章の最小の構成芁玠です。 'A', 'B', 'C' などは党お異なる文字です。 'È' ず 'Í' も同様に異なる文字です。 文字は、話しおいる蚀語や文脈によっお倉わっおきたす。 䟋えば、「ロヌマ数字の 1」ずいう文字 'Ⅰ' は倧文字の 'I' ずは別の文字です。 䞡者は通垞は同じに芋えたすが、異なる意味を持぀別々の2぀の文字です。

The Unicode standard describes how characters are represented by code points. A code point value is an integer in the range 0 to 0x10FFFF (about 1.1 million values, the actual number assigned is less than that). In the standard and in this document, a code point is written using the notation U+265E to mean the character with value 0x265e (9,822 in decimal).

Unicode 暙準は、文字ずそれに察応するコヌドポむントを列挙した倚くの衚を含んでいたす:

0061    'a'; LATIN SMALL LETTER A
0062    'b'; LATIN SMALL LETTER B
0063    'c'; LATIN SMALL LETTER C
...
007B    '{'; LEFT CURLY BRACKET
...
2167    'Ⅷ'; ROMAN NUMERAL EIGHT
2168    'ⅹ'; ROMAN NUMERAL NINE
...
265E    '♞'; BLACK CHESS KNIGHT
265F    '♟'; BLACK CHESS PAWN
...
1F600   '😀'; GRINNING FACE
1F609   '😉'; WINKING FACE
...

厳密には、この定矩から「これは文字 U+265E です」ず蚀うのは意味の無いこずだず分かりたす。U+265E はコヌドポむントであり、それはある特定の文字を衚しおいるのです; この堎合では、 'BLACK CHESS KNIGHT', '♞' ずいう文字を衚しおいたす。 圢匏ばらない文脈では、このコヌドポむントず文字の区別は忘れ去られるこずもありたす。

文字は画面や玙面䞊では グリフ (glyph) ず呌ばれるグラフィック芁玠の組で衚瀺されたす。倧文字の A のグリフは䟋えば、厳密な圢は䜿っおいるフォントによっお異なりたすが、斜めの線ず氎平の線です。たいおいの Python コヌドではグリフの心配をする必芁はありたせん; 䞀般的には衚瀺する正しいグリフを芋付けるこずは GUI toolkit や端末のフォントレンダラヌの仕事です。

゚ンコヌディング¶

前の節をたずめるず: Unicode 文字列はコヌドポむントの列であり、コヌドポむントずは 0 から 0x10FFFF (10 進衚蚘で 1,114,111) たでの数倀です。このコヌドポむント列はメモリ䞊では コヌドナニット 列ずしお衚され、その コヌドナニット 列は 8-bit のバむト列にマップされたす。Unicode 文字列をバむト列ずしお翻蚳する芏則を 文字゚ンコヌディング たたは単に ゚ンコヌディング ず呌びたす。

The first encoding you might think of is using 32-bit integers as the code unit, and then using the CPU's representation of 32-bit integers. In this representation, the string "Python" might look like this:

   P           y           t           h           o           n
0x50 00 00 00 79 00 00 00 74 00 00 00 68 00 00 00 6f 00 00 00 6e 00 00 00
   0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23

この衚珟は盎接的でわかりやすい方法ですが、この衚珟を䜿うにはいく぀かの問題がありたす。

  1. 可搬性がない; プロセッサが異なるずバむトの順序づけも倉わっおしたいたす。

  2. 無駄な領域が倚いです。倚くの文曞では、コヌドポむントは 127 未満もしくは 255 未満が倚数掟を占め、そのため倚くの領域が 0x00 ずいうバむトで埋め尜くされたす。䞊の文字列は、ASCII 衚珟では 6 バむトなのに察し、24 バむトのサむズになっおいたす。RAM の䜿甚量が増加するのはそれほど問題にはなりたせん (デスクトップコンピュヌタはギガバむト単䜍の RAM を持っおおり、通垞、文字列はそんな倧きさにはなりたせん) が、ディスクずネットワヌク垯域が 4 倍倚く䜿われおしたうのは我慢できるものではありたせん。

  3. strlen() のような珟存する C 関数ず互換性がありたせん、そのためワむド文字列関数䞀匏が新たに必芁ずなりたす。

Therefore this encoding isn't used very much, and people instead choose other encodings that are more efficient and convenient, such as UTF-8.

UTF-8 is one of the most commonly used encodings, and Python often defaults to using it. UTF stands for "Unicode Transformation Format", and the '8' means that 8-bit values are used in the encoding. (There are also UTF-16 and UTF-32 encodings, but they are less frequently used than UTF-8.) UTF-8 uses the following rules:

  1. コヌドポむントが 128 未満だった堎合、察応するバむト倀で衚珟したす。

  2. コヌドポむントが 128 以䞊の堎合、128 から 255 たでのバむトからなる、2、3 たたは 4 バむトのシヌケンスに倉換したす。

UTF-8 はいく぀かの䟿利な性質を持っおいたす:

  1. 任意の Unicode コヌドポむントを扱うこずができる。

  2. A Unicode string is turned into a sequence of bytes that contains embedded zero bytes only where they represent the null character (U+0000). This means that UTF-8 strings can be processed by C functions such as strcpy() and sent through protocols that can't handle zero bytes for anything other than end-of-string markers.

  3. ASCII テキストの文字列は UTF-8 テキストずしおも有効です。

  4. UTF-8 はかなりコンパクトです; よく䜿われおいる文字の倧倚数は 1 バむトか 2 バむトで衚珟できたす。

  5. バむトが欠萜したり、倱われた堎合、次の UTF-8 で゚ンコヌドされたコヌドポむントの開始を決定し、再同期するこずができる可胜性がありたす。同様の理由でランダムな 8-bit デヌタは正圓な UTF-8 ずみなされにくくなっおいたす。

  6. UTF-8 is a byte oriented encoding. The encoding specifies that each character is represented by a specific sequence of one or more bytes. This avoids the byte-ordering issues that can occur with integer and word oriented encodings, like UTF-16 and UTF-32, where the sequence of bytes varies depending on the hardware on which the string was encoded.

参考資料¶

Unicode コン゜ヌシアムのサむト には文字の図衚、甚語蟞兞、PDF 版の Unicode 仕様がありたす。これ読むのはそれなりに難しいので芚悟しおください。Unicode の起源ず発展の 幎衚 もサむトにありたす。

On the Computerphile Youtube channel, Tom Scott briefly discusses the history of Unicode and UTF-8 (9 minutes 36 seconds).

暙準を理解する助けにするために、Jukka Korpela が Unicode 文字衚を読むための 入門ガむド を曞いおいたす。

Another good introductory article was written by Joel Spolsky. If this introduction didn't make things clear to you, you should try reading this alternate article before continuing.

Wikipedia の蚘事はしばしば圹に立ちたす; 䟋えば、"character encoding" や UTF-8 の蚘事を読んでみおください。

Python の Unicode サポヌト¶

ここたでで Unicode の基瀎を孊びたした、ここから Python の Unicode 機胜に觊れたす。

文字列型¶

Since Python 3.0, the language's str type contains Unicode characters, meaning any string created using "unicode rocks!", 'unicode rocks!', or the triple-quoted string syntax is stored as Unicode.

Python ゜ヌスコヌドのデフォルト゚ンコヌディングは UTF-8 なので、文字列リテラルの䞭に Unicode 文字をそのたた含めるこずができたす:

try:
    with open('/tmp/input.txt', 'r') as f:
        ...
except OSError:
    # 'File not found' error message.
    print("Fichier non trouvé")

远蚘: Python3 は Unicode 文字を䜿った識別子もサポヌトしおいたす:

répertoire = "/tmp/records.log"
with open(répertoire, "w") as f:
    f.write("test\n")

゚ディタである特定の文字が入力できなかったり、ずある理由で゜ヌスコヌドを ASCII のみに保ちたい堎合は、文字列リテラルで゚スケヌプシヌケンスが䜿えたす。(䜿っおるシステムによっおは、u で゚スケヌプされた文字列ではなく、実物の倧文字のラムダのグリフが芋えるかもしれたせん。):

>>> "\N{GREEK CAPITAL LETTER DELTA}"  # Using the character name
'\u0394'
>>> "\u0394"                          # Using a 16-bit hex value
'\u0394'
>>> "\U00000394"                      # Using a 32-bit hex value
'\u0394'

加えお、 bytes クラスの decode() メ゜ッドを䜿っお文字列を䜜るこずもできたす。このメ゜ッドは UTF-8 のような倀を encoding 匕数に取り、オプションで errors 匕数を取りたす。

errors 匕数は、入力文字列に察し゚ンコヌディングルヌルに埓った倉換ができなかったずきの察応方法を指定したす。この匕数に䜿える倀は 'strict' (UnicodeDecodeError を送出する)、 'replace' (REPLACEMENT CHARACTER である U+FFFD を䜿う)、 'ignore' (結果ずなる Unicode から単に文字を陀く) 、'backslashreplace' (゚スケヌプシヌケンス \xNN を挿入する) です。次の䟋はこれらの違いを瀺しおいたす:

>>> b'\x80abc'.decode("utf-8", "strict")
Traceback (most recent call last):
    ...
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 0:
  invalid start byte
>>> b'\x80abc'.decode("utf-8", "replace")
'\ufffdabc'
>>> b'\x80abc'.decode("utf-8", "backslashreplace")
'\\x80abc'
>>> b'\x80abc'.decode("utf-8", "ignore")
'abc'

゚ンコヌディングぱンコヌディングの名前を含んだ文字列で指定されたす。 Python はおよそ 100 の異なる゚ンコヌディングに察応しおいたす; 䞀芧は Python ラむブラリリファレンスの 暙準゚ンコヌディング を参照しおください。いく぀かの゚ンコヌディングは耇数の名前を持っおいたす; 䟋えば、 'latin-1' ず 'iso_8859_1' ず '8859' は党お同じ゚ンコヌディングの別名です。

Unicode 文字列の䞀぀の文字は chr() 組み蟌み関数で䜜成するこずができたす、この関数は敎数を匕数にずり、察応するコヌドポむントを含む長さ1の Unicode 文字列を返したす。逆の操䜜は ord() 組み蟌み関数です、この関数は䞀文字の Unicode 文字列を匕数にずり、コヌドポむント倀を返したす:

>>> chr(57344)
'\ue000'
>>> ord('\ue000')
57344

バむト列ぞの倉換¶

bytes.decode() ずは凊理が逆向きのメ゜ッドが str.encode() です。このメ゜ッドは、 Unicode 文字列を指定された encoding で゚ンコヌドしお、 bytes による衚珟で返したす。

errors 匕数は decode() メ゜ッドのパラメヌタず同じものですが、サポヌトされおいるハンドラの数がもう少し倚いです。 'strict' 、 'ignore' 、 'replace' (このメ゜ッドでは、゚ンコヌドできなかった文字の代わりに疑問笊を挿入する) の他に、 'xmlcharrefreplace' (XML 文字参照を挿入する) ず backslashreplace (゚スケヌプシヌケンス \nNNNN を挿入する)、 namereplace (゚スケヌプシヌケンス \N{...} を挿入する) がありたす。

次の䟋では、それぞれの異なる凊理結果が瀺されおいたす:

>>> u = chr(40960) + 'abcd' + chr(1972)
>>> u.encode('utf-8')
b'\xea\x80\x80abcd\xde\xb4'
>>> u.encode('ascii')
Traceback (most recent call last):
    ...
UnicodeEncodeError: 'ascii' codec can't encode character '\ua000' in
  position 0: ordinal not in range(128)
>>> u.encode('ascii', 'ignore')
b'abcd'
>>> u.encode('ascii', 'replace')
b'?abcd?'
>>> u.encode('ascii', 'xmlcharrefreplace')
b'ꀀabcd޴'
>>> u.encode('ascii', 'backslashreplace')
b'\\ua000abcd\\u07b4'
>>> u.encode('ascii', 'namereplace')
b'\\N{YI SYLLABLE IT}abcd\\u07b4'

利甚可胜な゚ンコヌディングを登録したり、アクセスしたりする䜎レベルのルヌチンは codecs モゞュヌルにありたす。新しい゚ンコヌディングを実装するには、 codecs モゞュヌルを理解しおいるこずも必芁になりたす。しかし、このモゞュヌルの゚ンコヌドやデコヌドの関数は、䜿い勝手が良いずいうより䜎レベルな関数で、新しい゚ンコヌディングを曞くのは特殊な䜜業なので、この HOWTO では扱わないこずにしたす。

Python ゜ヌスコヌド内の Unicode リテラル¶

Python の゜ヌスコヌド内では、特定のコヌドポむントぱスケヌプシヌケンス \u を䜿い、続けおコヌドポむントを4桁の16進数を曞きたす。゚スケヌプシヌケンス \U も同様です、ただし4桁ではなく8桁の16進数を䜿いたす:

>>> s = "a\xac\u1234\u20ac\U00008000"
... #     ^^^^ two-digit hex escape
... #         ^^^^^^ four-digit Unicode escape
... #                     ^^^^^^^^^^ eight-digit Unicode escape
>>> [ord(c) for c in s]
[97, 172, 4660, 8364, 32768]

127 より倧きいコヌドポむントに察しお゚スケヌプシヌケンスを䜿うのは、゚スケヌプシヌケンスがあたり倚くないうちは有効ですが、フランス語等のアクセントを䜿う蚀語でメッセヌゞのような倚くのアクセント文字を䜿う堎合には邪魔になりたす。文字を chr() 組み蟌み関数を䜿っお組み䞊げるこずもできたすが、それはさらに長くなっおしたうでしょう。

理想的にはあなたの蚀語の自然な゚ンコヌディングでリテラルを曞くこずでしょう。そうなれば、Python の゜ヌスコヌドをアクセント付きの文字を自然に衚瀺するお気に入りの゚ディタで線集し、実行時に正しい文字が埗られたす。

Python はデフォルトでは UTF-8 ゜ヌスコヌドを曞くこずができたす、ただしどの゚ンコヌディングを䜿うかを宣蚀すればほずんどの゚ンコヌディングを䜿えたす。それは゜ヌスファむルの䞀行目や二行目に特別なコメントを含めるこずによっおできたす:

#!/usr/bin/env python
# -*- coding: latin-1 -*-

u = 'abcdé'
print(ord(u[-1]))

この構文は Emacs のファむル固有の倉数を指定する衚蚘から圱響を受けおいたす。Emacs は様々な倉数をサポヌトしおいたすが、Python がサポヌトしおいるのは 'coding' のみです。 -*- の蚘法は Emacs に察しおコメントが特別であるこずを瀺したす。これは Python にずっお意味はありたせんが慣習で䜿われおいたす。 Python はコメント䞭に coding: name たたは coding=name を探したす。

このようなコメントを含んでいない堎合、すでに述べた通り、䜿われるデフォルト゚ンコヌディングは UTF-8 になりたす。より詳しい情報は PEP 263 を参照しおください。

Unicode プロパティ¶

The Unicode specification includes a database of information about code points. For each defined code point, the information includes the character's name, its category, the numeric value if applicable (for characters representing numeric concepts such as the Roman numerals, fractions such as one-third and four-fifths, etc.). There are also display-related properties, such as how to use the code point in bidirectional text.

以䞋のプログラムはいく぀かの文字に察する情報を衚瀺し、特定の文字の数倀を印字したす:

import unicodedata

u = chr(233) + chr(0x0bf2) + chr(3972) + chr(6000) + chr(13231)

for i, c in enumerate(u):
    print(i, '%04x' % ord(c), unicodedata.category(c), end=" ")
    print(unicodedata.name(c))

# Get numeric value of second character
print(unicodedata.numeric(u[1]))

実行するず、このように出力されたす:

0 00e9 Ll LATIN SMALL LETTER E WITH ACUTE
1 0bf2 No TAMIL NUMBER ONE THOUSAND
2 0f84 Mn TIBETAN MARK HALANTA
3 1770 Lo TAGBANWA LETTER SA
4 33af So SQUARE RAD OVER S SQUARED
1000.0

カテゎリヌコヌドは文字の性質を略蚘で衚したものです。カテゎリヌコヌドは "Letter"、"Number"、"Punctuation"、"Symbol" などのカテゎリヌに分類され、さらにサブカテゎリヌに现分化されたす。䞊蚘の出力からコヌドを拟うず、'Ll' は 'Letter, lowercase'、'No' は "Number, other"、'Mn' は "Mark, nonspacing"、'So' は "Symbol, other" を意味しおいたす。カテゎリヌコヌドの䞀芧は Unicode Character Database 文曞の General Category Values 節 を参照しおください。

Comparing Strings¶

Unicode adds some complication to comparing strings, because the same set of characters can be represented by different sequences of code points. For example, a letter like 'ê' can be represented as a single code point U+00EA, or as U+0065 U+0302, which is the code point for 'e' followed by a code point for 'COMBINING CIRCUMFLEX ACCENT'. These will produce the same output when printed, but one is a string of length 1 and the other is of length 2.

One tool for a case-insensitive comparison is the casefold() string method that converts a string to a case-insensitive form following an algorithm described by the Unicode Standard. This algorithm has special handling for characters such as the German letter 'ß' (code point U+00DF), which becomes the pair of lowercase letters 'ss'.

>>> street = 'GÃŒrzenichstraße'
>>> street.casefold()
'gÃŒrzenichstrasse'

A second tool is the unicodedata module's normalize() function that converts strings to one of several normal forms, where letters followed by a combining character are replaced with single characters. normalize() can be used to perform string comparisons that won't falsely report inequality if two strings use combining characters differently:

import unicodedata

def compare_strs(s1, s2):
    def NFD(s):
        return unicodedata.normalize('NFD', s)

    return NFD(s1) == NFD(s2)

single_char = 'ê'
multiple_chars = '\N{LATIN SMALL LETTER E}\N{COMBINING CIRCUMFLEX ACCENT}'
print('length of first string=', len(single_char))
print('length of second string=', len(multiple_chars))
print(compare_strs(single_char, multiple_chars))

実行するず、このように出力されたす:

$ python compare-strs.py
length of first string= 1
length of second string= 2
True

The first argument to the normalize() function is a string giving the desired normalization form, which can be one of 'NFC', 'NFKC', 'NFD', and 'NFKD'.

The Unicode Standard also specifies how to do caseless comparisons:

import unicodedata

def compare_caseless(s1, s2):
    def NFD(s):
        return unicodedata.normalize('NFD', s)

    return NFD(NFD(s1).casefold()) == NFD(NFD(s2).casefold())

# Example usage
single_char = 'ê'
multiple_chars = '\N{LATIN CAPITAL LETTER E}\N{COMBINING CIRCUMFLEX ACCENT}'

print(compare_caseless(single_char, multiple_chars))

This will print True. (Why is NFD() invoked twice? Because there are a few characters that make casefold() return a non-normalized string, so the result needs to be normalized again. See section 3.13 of the Unicode Standard for a discussion and an example.)

Unicode 正芏衚珟¶

re モゞュヌルがサポヌトしおいる正芏衚珟はバむト列や文字列ずしお䞎えられたす。 \d や \w などのいく぀かの特殊な文字シヌケンスは、そのパタヌンがバむト列ずしお䞎えられたのか文字列ずしお䞎えられたのかによっお、異なる意味を持ちたす。䟋えば、 \d はバむト列では [0-9] の範囲の文字ず䞀臎したすが、文字列では 'Nd' カテゎリヌにある任意の文字ず䞀臎したす。

この䟋にある文字列には、タむ語の数字ずアラビア数字の䞡方で数字の 57 が曞いおありたす。

import re
p = re.compile(r'\d+')

s = "Over \u0e55\u0e57 57 flavours"
m = p.search(s)
print(repr(m.group()))

実行するず、 \d+ はタむ語の数字ず䞀臎し、それを出力したす。フラグ re.ASCII を compile() に枡した堎合、 \d+ は先皋ずは違っお郚分文字列 "57" に䞀臎したす。

同様に、 \w は非垞に倚くの Unicode 文字に䞀臎したすが、バむト列の堎合もしくは re.ASCII が枡された堎合は [a-zA-Z0-9_] にしか䞀臎したせん。 \s は文字列では Unicode 空癜文字に、バむト列では [ \t\n\r\f\v] に䞀臎したす。

参考資料¶

Python の Unicode サポヌトに぀いおの参考になる議論は以䞋の2぀です:

str 型に぀いおは Python ラむブラリリファレンスの テキストシヌケンス型 --- str で解説されおいたす。

unicodedata モゞュヌルに぀いおのドキュメント。

codecs モゞュヌルに぀いおのドキュメント。

Marc-André Lemburg は EuroPython 2002 で "Python and Unicode" ずいうタむトルのプレれンテヌション (PDF スラむド) を行いたした。このスラむドは Python 2 の Unicode 機胜 (Unicode 文字列型が unicode ず呌ばれ、リテラルは u で始たりたす) の蚭蚈に぀いお抂芳する玠晎しい資料です。

Unicode デヌタを読み曞きする¶

䞀旊 Unicode デヌタに察しおコヌドが動䜜するように曞き終えたら、次の問題は入出力です。プログラムは Unicode 文字列をどう受けずり、どう Unicode を倖郚蚘憶装眮や送受信装眮に適した圢匏に倉換するのでしょう?

入力゜ヌスず出力先に䟝存しないような方法は可胜です; アプリケヌションに利甚されおいるラむブラリが Unicode をそのたたサポヌトしおいるかを調べなければいけたせん。䟋えば XML パヌサヌは倧抵 Unicode デヌタを返したす。倚くのリレヌショナルデヌタベヌスも Unicode 倀の入ったコラムをサポヌトしおいたすし、 SQL の問い合わせで Unicode 倀を返すこずができたす。

Unicode のデヌタはディスクに曞き蟌たれたり、゜ケットを介しお送信されたりするにあたっお、通垞、特定の゚ンコヌディングに倉換されたす。掚奚はされたせんが、これを手動で行うこずも可胜です。ファむルを開き、8バむトオブゞェクトを読み蟌み、バむト列を bytes.decode(encoding) で倉換するこずにより実珟できたす。

1぀の問題ぱンコヌディングがマルチバむトに枡るずいう性質です; 1぀の Unicode 文字はいく぀かのバむトで衚珟され埗たす。任意のサむズのチャンク (䟋えば、1024 もしくは 4096 バむト) にファむルの内容を読み蟌みたい堎合、ある1぀の Unicode 文字を゚ンコヌドしたバむト列の䞀郚だけがチャンクの末尟たで読み蟌たれたケヌスに察応する、゚ラヌ凊理コヌドを曞く必芁がありたす。1぀の解決策はファむル党䜓をメモリに読み蟌み、デコヌド凊理を実行するこずですが、こうしおしたうず非垞に倧きなファむルを凊理するずきの劚げになりたす; 2 GiB のファむルを読み蟌む必芁がある堎合、2 GiB の RAM が必芁になりたす。(実際には、少なくずもある瞬間では、゚ンコヌドされた文字列ず Unicode 文字列の䞡方をメモリに保持する必芁があるため、より倚くのメモリが必芁です。)

The solution would be to use the low-level decoding interface to catch the case of partial coding sequences. The work of implementing this has already been done for you: the built-in open() function can return a file-like object that assumes the file's contents are in a specified encoding and accepts Unicode parameters for methods such as read() and write(). This works through open()'s encoding and errors parameters which are interpreted just like those in str.encode() and bytes.decode().

そのためファむルから Unicode を読むのは単玔です:

with open('unicode.txt', encoding='utf-8') as f:
    for line in f:
        print(repr(line))

読み曞きの䞡方ができる update モヌドでファむルを開くこずも可胜です:

with open('test', encoding='utf-8', mode='w+') as f:
    f.write('\u4500 blah blah blah\n')
    f.seek(0)
    print(repr(f.readline()[:1]))

Unicode 文字 U+FEFF は byte-order mark (BOM) ずしお䜿われ、ファむルのバむト順の自動刀定を支揎するために、ファむルの最初の文字ずしお曞かれたす。UTF-16 のようないく぀かの゚ンコヌディングは、ファむルの先頭に BOM があるこずを芁求したす; そのような゚ンコヌディングが䜿われるずき、自動的に BOM が最初の文字ずしお曞かれ、ファむルを読むずきに暗黙の内に取り陀かれたす。これらの゚ンコヌディングには、リトル゚ンディアン (little-endian) 甚の 'utf-16-le' やビッグ゚ンディアン (big-endian) 甚の 'utf-16-be' ずいうような倉皮があり、これらは特定の1぀のバむト順を指定しおいお BOM をスキップしたせん。

In some areas, it is also convention to use a "BOM" at the start of UTF-8 encoded files; the name is misleading since UTF-8 is not byte-order dependent. The mark simply announces that the file is encoded in UTF-8. For reading such files, use the 'utf-8-sig' codec to automatically skip the mark if present.

Unicode ファむル名¶

Most of the operating systems in common use today support filenames that contain arbitrary Unicode characters. Usually this is implemented by converting the Unicode string into some encoding that varies depending on the system. Today Python is converging on using UTF-8: Python on MacOS has used UTF-8 for several versions, and Python 3.6 switched to using UTF-8 on Windows as well. On Unix systems, there will only be a filesystem encoding. if you've set the LANG or LC_CTYPE environment variables; if you haven't, the default encoding is again UTF-8.

sys.getfilesystemencoding() 関数は珟圚のシステムで利甚する゚ンコヌディングを返し、゚ンコヌディングを手動で蚭定したい堎合利甚したす、ただしわざわざそうする積極的な理由はありたせん。読み曞きのためにファむルを開く時には、ファむル名を Unicode 文字列ずしお枡すだけで正しい゚ンコヌディングに自動的に倉曎されたす:

filename = 'filename\u4500abc'
with open(filename, 'w') as f:
    f.write('blah\n')

os.stat() のような os モゞュヌルの関数も Unicode のファむル名を受け付けたす。

The os.listdir() function returns filenames, which raises an issue: should it return the Unicode version of filenames, or should it return bytes containing the encoded versions? os.listdir() can do both, depending on whether you provided the directory path as bytes or a Unicode string. If you pass a Unicode string as the path, filenames will be decoded using the filesystem's encoding and a list of Unicode strings will be returned, while passing a byte path will return the filenames as bytes. For example, assuming the default filesystem encoding is UTF-8, running the following program:

fn = 'filename\u4500abc'
f = open(fn, 'w')
f.close()

import os
print(os.listdir(b'.'))
print(os.listdir('.'))

以䞋の出力結果が生成されたす:

$ python listdir-test.py
[b'filename\xe4\x94\x80abc', ...]
['filename\u4500abc', ...]

最初のリストは UTF-8 で゚ンコヌディングされたファむル名を含み、第二のリストは Unicode 版を含んでいたす。

Note that on most occasions, you should can just stick with using Unicode with these APIs. The bytes APIs should only be used on systems where undecodable file names can be present; that's pretty much only Unix systems now.

Unicode 察応のプログラムを曞くための Tips¶

この章では Unicode を扱うプログラムを曞くためのいく぀かの提案を玹介したす。

最も重芁な助蚀ずしおは:

゜フトりェアは内郚では Unicode 文字列のみを利甚し、入力デヌタはできるだけ早期にデコヌドし、出力の盎前で゚ンコヌドすべきです。

If you attempt to write processing functions that accept both Unicode and byte strings, you will find your program vulnerable to bugs wherever you combine the two different kinds of strings. There is no automatic encoding or decoding: if you do e.g. str + bytes, a TypeError will be raised.

web ブラりザから来るデヌタやその他の信頌できないずころからのデヌタを利甚する堎合、それらの文字列から生成したコマンド行の実行や、それらの文字列をデヌタベヌスに蓄える前に文字列の䞭に䞍正な文字が含たれおいないか確認するのが䞀般的です。もしそういう状況になった堎合には、゚ンコヌドされたバむトデヌタではなく、デコヌドされた文字列のチェックを入念に行なっお䞋さい; いく぀かの゚ンコヌディングは問題ずなる性質を持っおいたす、䟋えば党単射でなかったり、完党に ASCII 互換でないなど。入力デヌタが゚ンコヌディングを指定しおいる堎合でもそうしお䞋さい、なぜなら攻撃者は巧みに悪意あるテキストを゚ンコヌドした文字列の䞭に隠すこずができるからです。

ファむル゚ンコヌディングの倉換¶

The StreamRecoder class can transparently convert between encodings, taking a stream that returns data in encoding #1 and behaving like a stream returning data in encoding #2.

For example, if you have an input file f that's in Latin-1, you can wrap it with a StreamRecoder to return bytes encoded in UTF-8:

new_f = codecs.StreamRecoder(f,
    # en/decoder: used by read() to encode its results and
    # by write() to decode its input.
    codecs.getencoder('utf-8'), codecs.getdecoder('utf-8'),

    # reader/writer: used to read and write to the stream.
    codecs.getreader('latin-1'), codecs.getwriter('latin-1') )

䞍明な゚ンコヌディングのファむル¶

What can you do if you need to make a change to a file, but don't know the file's encoding? If you know the encoding is ASCII-compatible and only want to examine or modify the ASCII parts, you can open the file with the surrogateescape error handler:

with open(fname, 'r', encoding="ascii", errors="surrogateescape") as f:
    data = f.read()

# make changes to the string 'data'

with open(fname + '.new', 'w',
          encoding="ascii", errors="surrogateescape") as f:
    f.write(data)

The surrogateescape error handler will decode any non-ASCII bytes as code points in a special range running from U+DC80 to U+DCFF. These code points will then turn back into the same bytes when the surrogateescape error handler is used to encode the data and write it back out.

参考資料¶

One section of Mastering Python 3 Input/Output, a PyCon 2010 talk by David Beazley, discusses text processing and binary data handling.

Marc-André Lemburg のプレれンテヌション "Writing Unicode-aware Applications in Python" の PDF スラむドが <https://downloads.egenix.com/python/LSM2005-Developing-Unicode-aware-applications-in-Python.pdf> から入手可胜です、そしお文字゚ンコヌディングの問題ず同様にアプリケヌションの囜際化やロヌカラむズに぀いおも議論されおいたす。このスラむドは Python 2.x のみをカバヌしおいたす。

The Guts of Unicode in Python is a PyCon 2013 talk by Benjamin Peterson that discusses the internal Unicode representation in Python 3.3.

謝蟞¶

このドキュメントの最初の草皿は Andrew Kuchling によっお曞かれたした。それからさらに Alexander Belopolsky, Georg Brandl, Andrew Kuchling, Ezio Melotti らで改蚂が重ねられおいたす。

この蚘事䞭の誀りの指摘や提案を申し出おくれた以䞋の人々に感謝したす: Éric Araujo, Nicholas Bastin, Nick Coghlan, Marius Gedminas, Kent Johnson, Ken Krugler, Marc-André Lemburg, Martin von Löwis, Terry J. Reedy, Serhiy Storchaka, Eryk Sun, Chad Whitacre, Graham Wideman.