Unicode HOWTO¶
- Release:
1.12
ãã® HOWTO ææžã¯ãæåããŒã¿ã®è¡šçŸã®ããã® Unicode 仿§ã® Python ã«ããããµããŒãã«ã€ããŠè«ããããã« Unicode ã䜿ãããšãããšãã«ããåºå°ããå€ãã®åé¡ã«ã€ããŠèª¬æããŸãã
Unicode å ¥é¶
å®çŸ©Â¶
仿¥ã®ããã°ã©ã ã¯åºç¯å²ã®æåãæ±ããå¿ èŠããããŸãã ã¢ããªã±ãŒã·ã§ã³ã¯åœéåããããŠãŒã¶ãŒãéžã¹ãæ§ã ãªèšèªã§ã¡ãã»ãŒãžãåºåã衚瀺ããŸã; åãããã°ã©ã ããè±èªããã©ã³ã¹èªãæ¥æ¬èªãããã©ã€èªããã·ã¢èªã§ãšã©ãŒã¡ãã»ãŒãžãåºåããå¿ èŠããããããããŸããã Webã³ã³ãã³ãã¯ã©ããªèšèªã§ãæžãããå¯èœæ§ããããŸãããæ§ã ãªçµµæåãå«ãŸããããšããããŸãã Python ã®æåååã¯æå衚çŸã®ããã® Unicode æšæºã䜿ã£ãŠããŠã Python ããã°ã©ã ã¯æãåŸãæ§ã ãªæåãå šãŠæ±ããŸãã
Unicode (https://www.unicode.org/) ã¯ã人é¡ã®èšèªã§äœ¿ãããå šãŠã®æåãåæããããããã®æåèªèº«ã®äžæãªç¬Šå·ãäžããã®ãç®çãšãã仿§ã§ãã Unicode 仿§ã¯ç¶ç¶çã«æ¹èšãããæ°ããèšèªãèšå·ã远å ããæŽæ°ããªãããŠããŸãã
æå ã¯æç« ã®æå°ã®æ§æèŠçŽ ã§ãã 'A', 'B', 'C' ãªã©ã¯å šãŠç°ãªãæåã§ãã 'Ã' ãš 'Ã' ãåæ§ã«ç°ãªãæåã§ãã æåã¯ã話ããŠããèšèªãæèã«ãã£ãŠå€ãã£ãŠããŸãã äŸãã°ããããŒãæ°åã® 1ããšããæå 'â ' ã¯å€§æåã® 'I' ãšã¯å¥ã®æåã§ãã äž¡è ã¯éåžžã¯åãã«èŠããŸãããç°ãªãæå³ãæã€å¥ã ã®2ã€ã®æåã§ãã
The Unicode standard describes how characters are represented by
code points. A code point value is an integer in the range 0 to
0x10FFFF (about 1.1 million values, the
actual number assigned
is less than that). In the standard and in this document, a code point is written
using the notation U+265E to mean the character with value
0x265e (9,822 in decimal).
Unicode æšæºã¯ãæåãšããã«å¯Ÿå¿ããã³ãŒããã€ã³ããåæããå€ãã®è¡šãå«ãã§ããŸã:
0061 'a'; LATIN SMALL LETTER A
0062 'b'; LATIN SMALL LETTER B
0063 'c'; LATIN SMALL LETTER C
...
007B '{'; LEFT CURLY BRACKET
...
2167 'â
§'; ROMAN NUMERAL EIGHT
2168 'â
š'; ROMAN NUMERAL NINE
...
265E 'â'; BLACK CHESS KNIGHT
265F 'â'; BLACK CHESS PAWN
...
1F600 'ð'; GRINNING FACE
1F609 'ð'; WINKING FACE
...
å³å¯ã«ã¯ããã®å®çŸ©ãããããã¯æå U+265E ã§ãããšèšãã®ã¯æå³ã®ç¡ãããšã ãšåãããŸããU+265E ã¯ã³ãŒããã€ã³ãã§ãããããã¯ããç¹å®ã®æåã衚ããŠããã®ã§ã; ãã®å Žåã§ã¯ã 'BLACK CHESS KNIGHT', 'â' ãšããæåã衚ããŠããŸãã
圢åŒã°ããªãæèã§ã¯ããã®ã³ãŒããã€ã³ããšæåã®åºå¥ã¯å¿ãå»ãããããšããããŸãã
æåã¯ç»é¢ãçŽé¢äžã§ã¯ ã°ãªã (glyph) ãšåŒã°ããã°ã©ãã£ãã¯èŠçŽ ã®çµã§è¡šç€ºãããŸãã倧æåã® A ã®ã°ãªãã¯äŸãã°ãå³å¯ãªåœ¢ã¯äœ¿ã£ãŠãããã©ã³ãã«ãã£ãŠç°ãªããŸãããæãã®ç·ãšæ°Žå¹³ã®ç·ã§ãããããŠãã® Python ã³ãŒãã§ã¯ã°ãªãã®å¿é ãããå¿ èŠã¯ãããŸãã; äžè¬çã«ã¯è¡šç€ºããæ£ããã°ãªããèŠä»ããããšã¯ GUI toolkit ã端æ«ã®ãã©ã³ãã¬ã³ãã©ãŒã®ä»äºã§ãã
ãšã³ã³ãŒãã£ã³ã°Â¶
åã®ç¯ããŸãšãããš: Unicode æååã¯ã³ãŒããã€ã³ãã®åã§ãããã³ãŒããã€ã³ããšã¯ 0 ãã 0x10FFFF (10 é²è¡šèšã§ 1,114,111) ãŸã§ã®æ°å€ã§ãããã®ã³ãŒããã€ã³ãåã¯ã¡ã¢ãªäžã§ã¯ ã³ãŒããŠããã åãšããŠè¡šããããã® ã³ãŒããŠããã å㯠8-bit ã®ãã€ãåã«ããããããŸããUnicode æååããã€ãåãšããŠç¿»èš³ããèŠåã æåãšã³ã³ãŒãã£ã³ã° ãŸãã¯åã« ãšã³ã³ãŒãã£ã³ã° ãšåŒã³ãŸãã
The first encoding you might think of is using 32-bit integers as the code unit, and then using the CPU's representation of 32-bit integers. In this representation, the string "Python" might look like this:
P y t h o n
0x50 00 00 00 79 00 00 00 74 00 00 00 68 00 00 00 6f 00 00 00 6e 00 00 00
0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
ãã®è¡šçŸã¯çŽæ¥çã§ããããããæ¹æ³ã§ããããã®è¡šçŸã䜿ãã«ã¯ããã€ãã®åé¡ããããŸãã
坿¬æ§ããªã; ããã»ããµãç°ãªããšãã€ãã®é åºã¥ããå€ãã£ãŠããŸããŸãã
ç¡é§ãªé åãå€ãã§ããå€ãã®ææžã§ã¯ãã³ãŒããã€ã³ã㯠127 æªæºããã㯠255 æªæºã倿°æŽŸãå ãããã®ããå€ãã®é åã
0x00ãšãããã€ãã§åãå°œããããŸããäžã®æååã¯ãASCII 衚çŸã§ã¯ 6 ãã€ããªã®ã«å¯Ÿãã24 ãã€ãã®ãµã€ãºã«ãªã£ãŠããŸããRAM ã®äœ¿çšéãå¢å ããã®ã¯ããã»ã©åé¡ã«ã¯ãªããŸãã (ãã¹ã¯ãããã³ã³ãã¥ãŒã¿ã¯ã®ã¬ãã€ãåäœã® RAM ãæã£ãŠãããéåžžãæååã¯ãããªå€§ããã«ã¯ãªããŸãã) ãããã£ã¹ã¯ãšãããã¯ãŒã¯åž¯åã 4 åå€ã䜿ãããŠããŸãã®ã¯ææ ¢ã§ãããã®ã§ã¯ãããŸãããstrlen()ã®ãããªçŸåãã C 颿°ãšäºææ§ããããŸããããã®ããã¯ã€ãæåå颿°äžåŒãæ°ãã«å¿ èŠãšãªããŸãã
Therefore this encoding isn't used very much, and people instead choose other encodings that are more efficient and convenient, such as UTF-8.
UTF-8 is one of the most commonly used encodings, and Python often defaults to using it. UTF stands for "Unicode Transformation Format", and the '8' means that 8-bit values are used in the encoding. (There are also UTF-16 and UTF-32 encodings, but they are less frequently used than UTF-8.) UTF-8 uses the following rules:
ã³ãŒããã€ã³ãã 128 æªæºã ã£ãå Žåã察å¿ãããã€ãå€ã§è¡šçŸããŸãã
ã³ãŒããã€ã³ãã 128 以äžã®å Žåã128 ãã 255 ãŸã§ã®ãã€ããããªãã2ã3 ãŸã㯠4 ãã€ãã®ã·ãŒã±ã³ã¹ã«å€æããŸãã
UTF-8 ã¯ããã€ãã®äŸ¿å©ãªæ§è³ªãæã£ãŠããŸã:
ä»»æã® Unicode ã³ãŒããã€ã³ããæ±ãããšãã§ããã
A Unicode string is turned into a sequence of bytes that contains embedded zero bytes only where they represent the null character (U+0000). This means that UTF-8 strings can be processed by C functions such as
strcpy()and sent through protocols that can't handle zero bytes for anything other than end-of-string markers.ASCII ããã¹ãã®æåå㯠UTF-8 ããã¹ããšããŠãæå¹ã§ãã
UTF-8 ã¯ããªãã³ã³ãã¯ãã§ã; ãã䜿ãããŠããæåã®å€§å€æ°ã¯ 1 ãã€ãã 2 ãã€ãã§è¡šçŸã§ããŸãã
ãã€ããæ¬ èœãããã倱ãããå Žåãæ¬¡ã® UTF-8 ã§ãšã³ã³ãŒããããã³ãŒããã€ã³ãã®éå§ã決å®ããååæããããšãã§ããå¯èœæ§ããããŸããåæ§ã®çç±ã§ã©ã³ãã 㪠8-bit ããŒã¿ã¯æ£åœãª UTF-8 ãšã¿ãªããã«ãããªã£ãŠããŸãã
UTF-8 is a byte oriented encoding. The encoding specifies that each character is represented by a specific sequence of one or more bytes. This avoids the byte-ordering issues that can occur with integer and word oriented encodings, like UTF-16 and UTF-32, where the sequence of bytes varies depending on the hardware on which the string was encoded.
åèè³æÂ¶
Unicode ã³ã³ãœãŒã·ã¢ã ã®ãµã€ã ã«ã¯æåã®å³è¡šãçšèªèŸå žãPDF çã® Unicode 仿§ããããŸããããèªãã®ã¯ãããªãã«é£ããã®ã§èŠæããŠãã ãããUnicode ã®èµ·æºãšçºå±ã® 幎衚 ããµã€ãã«ãããŸãã
On the Computerphile Youtube channel, Tom Scott briefly discusses the history of Unicode and UTF-8 (9 minutes 36 seconds).
æšæºãçè§£ããå©ãã«ããããã«ãJukka Korpela ã Unicode æå衚ãèªãããã® å ¥éã¬ã€ã ãæžããŠããŸãã
Another good introductory article was written by Joel Spolsky. If this introduction didn't make things clear to you, you should try reading this alternate article before continuing.
Wikipedia ã®èšäºã¯ãã°ãã°åœ¹ã«ç«ã¡ãŸã; äŸãã°ã"character encoding" ã UTF-8 ã®èšäºãèªãã§ã¿ãŠãã ããã
Python ã® Unicode ãµããŒã¶
ãããŸã§ã§ Unicode ã®åºç€ãåŠã³ãŸããããããã Python ã® Unicode æ©èœã«è§ŠããŸãã
æååå¶
Since Python 3.0, the language's str type contains Unicode
characters, meaning any string created using "unicode rocks!", 'unicode
rocks!', or the triple-quoted string syntax is stored as Unicode.
Python ãœãŒã¹ã³ãŒãã®ããã©ã«ããšã³ã³ãŒãã£ã³ã°ã¯ UTF-8 ãªã®ã§ãæååãªãã©ã«ã®äžã« Unicode æåããã®ãŸãŸå«ããããšãã§ããŸã:
try:
with open('/tmp/input.txt', 'r') as f:
...
except OSError:
# 'File not found' error message.
print("Fichier non trouvé")
远èš: Python3 㯠Unicode æåã䜿ã£ãèå¥åããµããŒãããŠããŸã:
répertoire = "/tmp/records.log"
with open(répertoire, "w") as f:
f.write("test\n")
ãšãã£ã¿ã§ããç¹å®ã®æåãå ¥åã§ããªãã£ããããšããçç±ã§ãœãŒã¹ã³ãŒãã ASCII ã®ã¿ã«ä¿ã¡ããå Žåã¯ãæååãªãã©ã«ã§ãšã¹ã±ãŒãã·ãŒã±ã³ã¹ã䜿ããŸãã(䜿ã£ãŠãã·ã¹ãã ã«ãã£ãŠã¯ãu ã§ãšã¹ã±ãŒããããæååã§ã¯ãªããå®ç©ã®å€§æåã®ã©ã ãã®ã°ãªããèŠãããããããŸããã):
>>> "\N{GREEK CAPITAL LETTER DELTA}" # Using the character name
'\u0394'
>>> "\u0394" # Using a 16-bit hex value
'\u0394'
>>> "\U00000394" # Using a 32-bit hex value
'\u0394'
å ããŠã bytes ã¯ã©ã¹ã® decode() ã¡ãœããã䜿ã£ãŠæååãäœãããšãã§ããŸãããã®ã¡ãœãã㯠UTF-8 ã®ãããªå€ã encoding åŒæ°ã«åãããªãã·ã§ã³ã§ errors åŒæ°ãåããŸãã
errors åŒæ°ã¯ãå
¥åæååã«å¯Ÿããšã³ã³ãŒãã£ã³ã°ã«ãŒã«ã«åŸã£ã倿ãã§ããªãã£ããšãã®å¯Ÿå¿æ¹æ³ãæå®ããŸãããã®åŒæ°ã«äœ¿ããå€ã¯ 'strict' (UnicodeDecodeError ãéåºãã)ã 'replace' (REPLACEMENT CHARACTER ã§ãã U+FFFD ã䜿ã)ã 'ignore' (çµæãšãªã Unicode ããåã«æåãé€ã) ã'backslashreplace' (ãšã¹ã±ãŒãã·ãŒã±ã³ã¹ \xNN ãæ¿å
¥ãã) ã§ããæ¬¡ã®äŸã¯ãããã®éãã瀺ããŠããŸã:
>>> b'\x80abc'.decode("utf-8", "strict")
Traceback (most recent call last):
...
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 0:
invalid start byte
>>> b'\x80abc'.decode("utf-8", "replace")
'\ufffdabc'
>>> b'\x80abc'.decode("utf-8", "backslashreplace")
'\\x80abc'
>>> b'\x80abc'.decode("utf-8", "ignore")
'abc'
ãšã³ã³ãŒãã£ã³ã°ã¯ãšã³ã³ãŒãã£ã³ã°ã®ååãå«ãã æååã§æå®ãããŸãã Python ã¯ããã 100 ã®ç°ãªããšã³ã³ãŒãã£ã³ã°ã«å¯Ÿå¿ããŠããŸã; äžèŠ§ã¯ Python ã©ã€ãã©ãªãªãã¡ã¬ã³ã¹ã® æšæºãšã³ã³ãŒãã£ã³ã° ãåç
§ããŠãã ãããããã€ãã®ãšã³ã³ãŒãã£ã³ã°ã¯è€æ°ã®ååãæã£ãŠããŸã; äŸãã°ã 'latin-1' ãš 'iso_8859_1' ãš '8859' ã¯å
šãŠåããšã³ã³ãŒãã£ã³ã°ã®å¥åã§ãã
Unicode æååã®äžã€ã®æå㯠chr() çµã¿èŸŒã¿é¢æ°ã§äœæããããšãã§ããŸãããã®é¢æ°ã¯æŽæ°ãåŒæ°ã«ãšãã察å¿ããã³ãŒããã€ã³ããå«ãé·ã1ã® Unicode æååãè¿ããŸããéã®æäœã¯ ord() çµã¿èŸŒã¿é¢æ°ã§ãããã®é¢æ°ã¯äžæåã® Unicode æååãåŒæ°ã«ãšããã³ãŒããã€ã³ãå€ãè¿ããŸã:
>>> chr(57344)
'\ue000'
>>> ord('\ue000')
57344
ãã€ãåãžã®å€æÂ¶
bytes.decode() ãšã¯åŠçãéåãã®ã¡ãœããã str.encode() ã§ãããã®ã¡ãœããã¯ã Unicode æååãæå®ããã encoding ã§ãšã³ã³ãŒãããŠã bytes ã«ãã衚çŸã§è¿ããŸãã
errors åŒæ°ã¯ decode() ã¡ãœããã®ãã©ã¡ãŒã¿ãšåããã®ã§ããããµããŒããããŠãããã³ãã©ã®æ°ãããå°ãå€ãã§ãã
'strict' ã 'ignore' ã 'replace' (ãã®ã¡ãœããã§ã¯ããšã³ã³ãŒãã§ããªãã£ãæåã®ä»£ããã«çåç¬Šãæ¿å
¥ãã) ã®ä»ã«ã 'xmlcharrefreplace' (XML æååç
§ãæ¿å
¥ãã) ãš backslashreplace (ãšã¹ã±ãŒãã·ãŒã±ã³ã¹ \nNNNN ãæ¿å
¥ãã)ã namereplace (ãšã¹ã±ãŒãã·ãŒã±ã³ã¹ \N{...} ãæ¿å
¥ãã) ããããŸãã
次ã®äŸã§ã¯ãããããã®ç°ãªãåŠççµæã瀺ãããŠããŸã:
>>> u = chr(40960) + 'abcd' + chr(1972)
>>> u.encode('utf-8')
b'\xea\x80\x80abcd\xde\xb4'
>>> u.encode('ascii')
Traceback (most recent call last):
...
UnicodeEncodeError: 'ascii' codec can't encode character '\ua000' in
position 0: ordinal not in range(128)
>>> u.encode('ascii', 'ignore')
b'abcd'
>>> u.encode('ascii', 'replace')
b'?abcd?'
>>> u.encode('ascii', 'xmlcharrefreplace')
b'ꀀabcd޴'
>>> u.encode('ascii', 'backslashreplace')
b'\\ua000abcd\\u07b4'
>>> u.encode('ascii', 'namereplace')
b'\\N{YI SYLLABLE IT}abcd\\u07b4'
å©çšå¯èœãªãšã³ã³ãŒãã£ã³ã°ãç»é²ããããã¢ã¯ã»ã¹ãããããäœã¬ãã«ã®ã«ãŒãã³ã¯ codecs ã¢ãžã¥ãŒã«ã«ãããŸããæ°ãããšã³ã³ãŒãã£ã³ã°ãå®è£
ããã«ã¯ã codecs ã¢ãžã¥ãŒã«ãçè§£ããŠããããšãå¿
èŠã«ãªããŸãããããããã®ã¢ãžã¥ãŒã«ã®ãšã³ã³ãŒãããã³ãŒãã®é¢æ°ã¯ã䜿ãåæãè¯ããšããããäœã¬ãã«ãªé¢æ°ã§ãæ°ãããšã³ã³ãŒãã£ã³ã°ãæžãã®ã¯ç¹æ®ãªäœæ¥ãªã®ã§ããã® HOWTO ã§ã¯æ±ããªãããšã«ããŸãã
Python ãœãŒã¹ã³ãŒãå ã® Unicode ãªãã©ã«Â¶
Python ã®ãœãŒã¹ã³ãŒãå
ã§ã¯ãç¹å®ã®ã³ãŒããã€ã³ãã¯ãšã¹ã±ãŒãã·ãŒã±ã³ã¹ \u ã䜿ããç¶ããŠã³ãŒããã€ã³ãã4æ¡ã®16鲿°ãæžããŸãããšã¹ã±ãŒãã·ãŒã±ã³ã¹ \U ãåæ§ã§ãããã ã4æ¡ã§ã¯ãªã8æ¡ã®16鲿°ã䜿ããŸã:
>>> s = "a\xac\u1234\u20ac\U00008000"
... # ^^^^ two-digit hex escape
... # ^^^^^^ four-digit Unicode escape
... # ^^^^^^^^^^ eight-digit Unicode escape
>>> [ord(c) for c in s]
[97, 172, 4660, 8364, 32768]
127 ãã倧ããã³ãŒããã€ã³ãã«å¯ŸããŠãšã¹ã±ãŒãã·ãŒã±ã³ã¹ã䜿ãã®ã¯ããšã¹ã±ãŒãã·ãŒã±ã³ã¹ãããŸãå€ããªããã¡ã¯æå¹ã§ããããã©ã³ã¹èªçã®ã¢ã¯ã»ã³ãã䜿ãèšèªã§ã¡ãã»ãŒãžã®ãããªå€ãã®ã¢ã¯ã»ã³ãæåã䜿ãå Žåã«ã¯éªéã«ãªããŸããæåã chr() çµã¿èŸŒã¿é¢æ°ã䜿ã£ãŠçµã¿äžããããšãã§ããŸãããããã¯ããã«é·ããªã£ãŠããŸãã§ãããã
çæ³çã«ã¯ããªãã®èšèªã®èªç¶ãªãšã³ã³ãŒãã£ã³ã°ã§ãªãã©ã«ãæžãããšã§ãããããããªãã°ãPython ã®ãœãŒã¹ã³ãŒããã¢ã¯ã»ã³ãä»ãã®æåãèªç¶ã«è¡šç€ºãããæ°ã«å ¥ãã®ãšãã£ã¿ã§ç·šéããå®è¡æã«æ£ããæåãåŸãããŸãã
Python ã¯ããã©ã«ãã§ã¯ UTF-8 ãœãŒã¹ã³ãŒããæžãããšãã§ããŸãããã ãã©ã®ãšã³ã³ãŒãã£ã³ã°ã䜿ããã宣èšããã°ã»ãšãã©ã®ãšã³ã³ãŒãã£ã³ã°ã䜿ããŸããããã¯ãœãŒã¹ãã¡ã€ã«ã®äžè¡ç®ãäºè¡ç®ã«ç¹å¥ãªã³ã¡ã³ããå«ããããšã«ãã£ãŠã§ããŸã:
#!/usr/bin/env python
# -*- coding: latin-1 -*-
u = 'abcdé'
print(ord(u[-1]))
ãã®æ§æã¯ Emacs ã®ãã¡ã€ã«åºæã®å€æ°ãæå®ãã衚èšãã圱é¿ãåããŠããŸããEmacs ã¯æ§ã
ãªå€æ°ããµããŒãããŠããŸãããPython ããµããŒãããŠããã®ã¯ 'coding' ã®ã¿ã§ãã -*- ã®èšæ³ã¯ Emacs ã«å¯ŸããŠã³ã¡ã³ããç¹å¥ã§ããããšã瀺ããŸãããã㯠Python ã«ãšã£ãŠæå³ã¯ãããŸãããæ
£ç¿ã§äœ¿ãããŠããŸãã Python ã¯ã³ã¡ã³ãäžã« coding: name ãŸã㯠coding=name ãæ¢ããŸãã
ãã®ãããªã³ã¡ã³ããå«ãã§ããªãå Žåããã§ã«è¿°ã¹ãéãã䜿ãããããã©ã«ããšã³ã³ãŒãã£ã³ã°ã¯ UTF-8 ã«ãªããŸãããã詳ããæ å ±ã¯ PEP 263 ãåç §ããŠãã ããã
Unicode ããããã£Â¶
The Unicode specification includes a database of information about code points. For each defined code point, the information includes the character's name, its category, the numeric value if applicable (for characters representing numeric concepts such as the Roman numerals, fractions such as one-third and four-fifths, etc.). There are also display-related properties, such as how to use the code point in bidirectional text.
以äžã®ããã°ã©ã ã¯ããã€ãã®æåã«å¯Ÿããæ å ±ã衚瀺ããç¹å®ã®æåã®æ°å€ãå°åããŸã:
import unicodedata
u = chr(233) + chr(0x0bf2) + chr(3972) + chr(6000) + chr(13231)
for i, c in enumerate(u):
print(i, '%04x' % ord(c), unicodedata.category(c), end=" ")
print(unicodedata.name(c))
# Get numeric value of second character
print(unicodedata.numeric(u[1]))
å®è¡ãããšããã®ããã«åºåãããŸã:
0 00e9 Ll LATIN SMALL LETTER E WITH ACUTE
1 0bf2 No TAMIL NUMBER ONE THOUSAND
2 0f84 Mn TIBETAN MARK HALANTA
3 1770 Lo TAGBANWA LETTER SA
4 33af So SQUARE RAD OVER S SQUARED
1000.0
ã«ããŽãªãŒã³ãŒãã¯æåã®æ§è³ªãç¥èšã§è¡šãããã®ã§ããã«ããŽãªãŒã³ãŒã㯠"Letter"ã"Number"ã"Punctuation"ã"Symbol" ãªã©ã®ã«ããŽãªãŒã«åé¡ãããããã«ãµãã«ããŽãªãŒã«çްååãããŸããäžèšã®åºåããã³ãŒããæŸããšã'Ll' 㯠'Letter, lowercase'ã'No' 㯠"Number, other"ã'Mn' 㯠"Mark, nonspacing"ã'So' 㯠"Symbol, other" ãæå³ããŠããŸããã«ããŽãªãŒã³ãŒãã®äžèŠ§ã¯ Unicode Character Database ææžã® General Category Values ç¯ ãåç
§ããŠãã ããã
Comparing Strings¶
Unicode adds some complication to comparing strings, because the same set of characters can be represented by different sequences of code points. For example, a letter like 'ê' can be represented as a single code point U+00EA, or as U+0065 U+0302, which is the code point for 'e' followed by a code point for 'COMBINING CIRCUMFLEX ACCENT'. These will produce the same output when printed, but one is a string of length 1 and the other is of length 2.
One tool for a case-insensitive comparison is the
casefold() string method that converts a string to a
case-insensitive form following an algorithm described by the Unicode
Standard. This algorithm has special handling for characters such as
the German letter 'Ã' (code point U+00DF), which becomes the pair of
lowercase letters 'ss'.
>>> street = 'GÃŒrzenichstraÃe'
>>> street.casefold()
'gÃŒrzenichstrasse'
A second tool is the unicodedata module's
normalize() function that converts strings to one
of several normal forms, where letters followed by a combining character are
replaced with single characters. normalize() can
be used to perform string comparisons that won't falsely report
inequality if two strings use combining characters differently:
import unicodedata
def compare_strs(s1, s2):
def NFD(s):
return unicodedata.normalize('NFD', s)
return NFD(s1) == NFD(s2)
single_char = 'ê'
multiple_chars = '\N{LATIN SMALL LETTER E}\N{COMBINING CIRCUMFLEX ACCENT}'
print('length of first string=', len(single_char))
print('length of second string=', len(multiple_chars))
print(compare_strs(single_char, multiple_chars))
å®è¡ãããšããã®ããã«åºåãããŸã:
$ python compare-strs.py
length of first string= 1
length of second string= 2
True
The first argument to the normalize() function is a
string giving the desired normalization form, which can be one of
'NFC', 'NFKC', 'NFD', and 'NFKD'.
The Unicode Standard also specifies how to do caseless comparisons:
import unicodedata
def compare_caseless(s1, s2):
def NFD(s):
return unicodedata.normalize('NFD', s)
return NFD(NFD(s1).casefold()) == NFD(NFD(s2).casefold())
# Example usage
single_char = 'ê'
multiple_chars = '\N{LATIN CAPITAL LETTER E}\N{COMBINING CIRCUMFLEX ACCENT}'
print(compare_caseless(single_char, multiple_chars))
This will print True. (Why is NFD() invoked twice? Because
there are a few characters that make casefold() return a
non-normalized string, so the result needs to be normalized again. See
section 3.13 of the Unicode Standard for a discussion and an example.)
Unicode æ£èŠè¡šçŸÂ¶
re ã¢ãžã¥ãŒã«ããµããŒãããŠããæ£èŠè¡šçŸã¯ãã€ãåãæååãšããŠäžããããŸãã \d ã \w ãªã©ã®ããã€ãã®ç¹æ®ãªæåã·ãŒã±ã³ã¹ã¯ããã®ãã¿ãŒã³ããã€ãåãšããŠäžããããã®ãæååãšããŠäžããããã®ãã«ãã£ãŠãç°ãªãæå³ãæã¡ãŸããäŸãã°ã \d ã¯ãã€ãåã§ã¯ [0-9] ã®ç¯å²ã®æåãšäžèŽããŸãããæååã§ã¯ 'Nd' ã«ããŽãªãŒã«ããä»»æã®æåãšäžèŽããŸãã
ãã®äŸã«ããæååã«ã¯ãã¿ã€èªã®æ°åãšã¢ã©ãã¢æ°åã®äž¡æ¹ã§æ°åã® 57 ãæžããŠãããŸãã
import re
p = re.compile(r'\d+')
s = "Over \u0e55\u0e57 57 flavours"
m = p.search(s)
print(repr(m.group()))
å®è¡ãããšã \d+ ã¯ã¿ã€èªã®æ°åãšäžèŽãããããåºåããŸãããã©ã° re.ASCII ã compile() ã«æž¡ããå Žåã \d+ ã¯å
çšãšã¯éã£ãŠéšåæåå "57" ã«äžèŽããŸãã
åæ§ã«ã \w ã¯éåžžã«å€ãã® Unicode æåã«äžèŽããŸããããã€ãåã®å Žåããã㯠re.ASCII ãæž¡ãããå Žå㯠[a-zA-Z0-9_] ã«ããäžèŽããŸããã \s ã¯æååã§ã¯ Unicode ç©ºçœæåã«ããã€ãåã§ã¯ [ \t\n\r\f\v] ã«äžèŽããŸãã
åèè³æÂ¶
Python ã® Unicode ãµããŒãã«ã€ããŠã®åèã«ãªãè°è«ã¯ä»¥äžã®2ã€ã§ã:
Nick Coghlan ã«ãã Processing Text Files in Python 3
Ned Batchelder ã«ãã PyCon 2012 ã§ã®çºè¡š Pragmatic Unicode
str åã«ã€ããŠã¯ Python ã©ã€ãã©ãªãªãã¡ã¬ã³ã¹ã® ããã¹ãã·ãŒã±ã³ã¹å --- str ã§è§£èª¬ãããŠããŸãã
unicodedata ã¢ãžã¥ãŒã«ã«ã€ããŠã®ããã¥ã¡ã³ãã
codecs ã¢ãžã¥ãŒã«ã«ã€ããŠã®ããã¥ã¡ã³ãã
Marc-André Lemburg 㯠EuroPython 2002 ã§ "Python and Unicode" ãšããã¿ã€ãã«ã®ãã¬ãŒã³ããŒã·ã§ã³ (PDF ã¹ã©ã€ã) ãè¡ããŸããããã®ã¹ã©ã€ã㯠Python 2 ã® Unicode æ©èœ (Unicode æåååã unicode ãšåŒã°ãããªãã©ã«ã¯ u ã§å§ãŸããŸã) ã®èšèšã«ã€ããŠæŠèгããçŽ æŽããè³æã§ãã
Unicode ããŒã¿ãèªã¿æžããã¶
äžæŠ Unicode ããŒã¿ã«å¯ŸããŠã³ãŒããåäœããããã«æžãçµããããæ¬¡ã®åé¡ã¯å ¥åºåã§ããããã°ã©ã 㯠Unicode æååãã©ãåããšããã©ã Unicode ãå€éšèšæ¶è£ 眮ãéåä¿¡è£ çœ®ã«é©ãã圢åŒã«å€æããã®ã§ããã?
å ¥åãœãŒã¹ãšåºåå ã«äŸåããªããããªæ¹æ³ã¯å¯èœã§ã; ã¢ããªã±ãŒã·ã§ã³ã«å©çšãããŠããã©ã€ãã©ãªã Unicode ããã®ãŸãŸãµããŒãããŠãããã調ã¹ãªããã°ãããŸãããäŸãã° XML ããŒãµãŒã¯å€§æµ Unicode ããŒã¿ãè¿ããŸããå€ãã®ãªã¬ãŒã·ã§ãã«ããŒã¿ããŒã¹ã Unicode å€ã®å ¥ã£ãã³ã©ã ããµããŒãããŠããŸããã SQL ã®åãåããã§ Unicode å€ãè¿ãããšãã§ããŸãã
Unicode ã®ããŒã¿ã¯ãã£ã¹ã¯ã«æžã蟌ãŸãããããœã±ãããä»ããŠéä¿¡ããããããã«ããã£ãŠãéåžžãç¹å®ã®ãšã³ã³ãŒãã£ã³ã°ã«å€æãããŸããæšå¥šã¯ãããŸãããããããæåã§è¡ãããšãå¯èœã§ãããã¡ã€ã«ãéãã8ãã€ããªããžã§ã¯ããèªã¿èŸŒã¿ããã€ãåã bytes.decode(encoding) ã§å€æããããšã«ããå®çŸã§ããŸãã
1ã€ã®åé¡ã¯ãšã³ã³ãŒãã£ã³ã°ããã«ããã€ãã«æž¡ããšããæ§è³ªã§ã; 1ã€ã® Unicode æåã¯ããã€ãã®ãã€ãã§è¡šçŸããåŸãŸããä»»æã®ãµã€ãºã®ãã£ã³ã¯ (äŸãã°ã1024 ããã㯠4096 ãã€ã) ã«ãã¡ã€ã«ã®å 容ãèªã¿èŸŒã¿ããå Žåããã1ã€ã® Unicode æåããšã³ã³ãŒããããã€ãåã®äžéšã ãããã£ã³ã¯ã®æ«å°ŸãŸã§èªã¿èŸŒãŸããã±ãŒã¹ã«å¯Ÿå¿ããããšã©ãŒåŠçã³ãŒããæžãå¿ èŠããããŸãã1ã€ã®è§£æ±ºçã¯ãã¡ã€ã«å šäœãã¡ã¢ãªã«èªã¿èŸŒã¿ããã³ãŒãåŠçãå®è¡ããããšã§ãããããããŠããŸããšéåžžã«å€§ããªãã¡ã€ã«ãåŠçãããšãã®åŠšãã«ãªããŸã; 2 GiB ã®ãã¡ã€ã«ãèªã¿èŸŒãå¿ èŠãããå Žåã2 GiB ã® RAM ãå¿ èŠã«ãªããŸãã(å®éã«ã¯ãå°ãªããšãããç¬éã§ã¯ããšã³ã³ãŒããããæååãš Unicode æååã®äž¡æ¹ãã¡ã¢ãªã«ä¿æããå¿ èŠããããããããå€ãã®ã¡ã¢ãªãå¿ èŠã§ãã)
The solution would be to use the low-level decoding interface to catch the case
of partial coding sequences. The work of implementing this has already been
done for you: the built-in open() function can return a file-like object
that assumes the file's contents are in a specified encoding and accepts Unicode
parameters for methods such as read() and
write(). This works through open()'s encoding and
errors parameters which are interpreted just like those in str.encode()
and bytes.decode().
ãã®ãããã¡ã€ã«ãã Unicode ãèªãã®ã¯åçŽã§ã:
with open('unicode.txt', encoding='utf-8') as f:
for line in f:
print(repr(line))
èªã¿æžãã®äž¡æ¹ãã§ãã update ã¢ãŒãã§ãã¡ã€ã«ãéãããšãå¯èœã§ã:
with open('test', encoding='utf-8', mode='w+') as f:
f.write('\u4500 blah blah blah\n')
f.seek(0)
print(repr(f.readline()[:1]))
Unicode æå U+FEFF 㯠byte-order mark (BOM) ãšããŠäœ¿ããããã¡ã€ã«ã®ãã€ãé ã®èªåå€å®ãæ¯æŽããããã«ããã¡ã€ã«ã®æåã®æåãšããŠæžãããŸããUTF-16 ã®ãããªããã€ãã®ãšã³ã³ãŒãã£ã³ã°ã¯ããã¡ã€ã«ã®å
é ã« BOM ãããããšãèŠæ±ããŸã; ãã®ãããªãšã³ã³ãŒãã£ã³ã°ã䜿ããããšããèªåçã« BOM ãæåã®æåãšããŠæžããããã¡ã€ã«ãèªããšãã«æé»ã®å
ã«åãé€ãããŸãããããã®ãšã³ã³ãŒãã£ã³ã°ã«ã¯ããªãã«ãšã³ãã£ã¢ã³ (little-endian) çšã® 'utf-16-le' ãããã°ãšã³ãã£ã¢ã³ (big-endian) çšã® 'utf-16-be' ãšãããããªå€çš®ãããããããã¯ç¹å®ã®1ã€ã®ãã€ãé ãæå®ããŠã㊠BOM ãã¹ãããããŸããã
In some areas, it is also convention to use a "BOM" at the start of UTF-8 encoded files; the name is misleading since UTF-8 is not byte-order dependent. The mark simply announces that the file is encoded in UTF-8. For reading such files, use the 'utf-8-sig' codec to automatically skip the mark if present.
Unicode ãã¡ã€ã«å¶
Most of the operating systems in common use today support filenames
that contain arbitrary Unicode characters. Usually this is
implemented by converting the Unicode string into some encoding that
varies depending on the system. Today Python is converging on using
UTF-8: Python on MacOS has used UTF-8 for several versions, and Python
3.6 switched to using UTF-8 on Windows as well. On Unix systems,
there will only be a filesystem encoding. if you've set the LANG or LC_CTYPE environment variables; if
you haven't, the default encoding is again UTF-8.
sys.getfilesystemencoding() 颿°ã¯çŸåšã®ã·ã¹ãã ã§å©çšãããšã³ã³ãŒãã£ã³ã°ãè¿ãããšã³ã³ãŒãã£ã³ã°ãæåã§èšå®ãããå Žåå©çšããŸãããã ãããããããããç©æ¥µçãªçç±ã¯ãããŸãããèªã¿æžãã®ããã«ãã¡ã€ã«ãéãæã«ã¯ããã¡ã€ã«åã Unicode æååãšããŠæž¡ãã ãã§æ£ãããšã³ã³ãŒãã£ã³ã°ã«èªåçã«å€æŽãããŸã:
filename = 'filename\u4500abc'
with open(filename, 'w') as f:
f.write('blah\n')
os.stat() ã®ãã㪠os ã¢ãžã¥ãŒã«ã®é¢æ°ã Unicode ã®ãã¡ã€ã«åãåãä»ããŸãã
The os.listdir() function returns filenames, which raises an issue: should it return
the Unicode version of filenames, or should it return bytes containing
the encoded versions? os.listdir() can do both, depending on whether you
provided the directory path as bytes or a Unicode string. If you pass a
Unicode string as the path, filenames will be decoded using the filesystem's
encoding and a list of Unicode strings will be returned, while passing a byte
path will return the filenames as bytes. For example,
assuming the default filesystem encoding is UTF-8, running the following program:
fn = 'filename\u4500abc'
f = open(fn, 'w')
f.close()
import os
print(os.listdir(b'.'))
print(os.listdir('.'))
以äžã®åºåçµæãçæãããŸã:
$ python listdir-test.py
[b'filename\xe4\x94\x80abc', ...]
['filename\u4500abc', ...]
æåã®ãªã¹ã㯠UTF-8 ã§ãšã³ã³ãŒãã£ã³ã°ããããã¡ã€ã«åãå«ã¿ã第äºã®ãªã¹ã㯠Unicode çãå«ãã§ããŸãã
Note that on most occasions, you should can just stick with using Unicode with these APIs. The bytes APIs should only be used on systems where undecodable file names can be present; that's pretty much only Unix systems now.
Unicode 察å¿ã®ããã°ã©ã ãæžãããã® Tips¶
ãã®ç« ã§ã¯ Unicode ãæ±ãããã°ã©ã ãæžãããã®ããã€ãã®ææ¡ã玹ä»ããŸãã
æãéèŠãªå©èšãšããŠã¯:
ãœãããŠã§ã¢ã¯å éšã§ã¯ Unicode æååã®ã¿ãå©çšããå ¥åããŒã¿ã¯ã§ããã ãæ©æã«ãã³ãŒãããåºåã®çŽåã§ãšã³ã³ãŒããã¹ãã§ãã
If you attempt to write processing functions that accept both Unicode and byte
strings, you will find your program vulnerable to bugs wherever you combine the
two different kinds of strings. There is no automatic encoding or decoding: if
you do e.g. str + bytes, a TypeError will be raised.
web ãã©ãŠã¶ããæ¥ãããŒã¿ããã®ä»ã®ä¿¡é Œã§ããªããšããããã®ããŒã¿ãå©çšããå Žåããããã®æååããçæããã³ãã³ãè¡ã®å®è¡ãããããã®æååãããŒã¿ããŒã¹ã«èããåã«æååã®äžã«äžæ£ãªæåãå«ãŸããŠããªãã確èªããã®ãäžè¬çã§ããããããããç¶æ³ã«ãªã£ãå Žåã«ã¯ããšã³ã³ãŒãããããã€ãããŒã¿ã§ã¯ãªãããã³ãŒããããæååã®ãã§ãã¯ãå ¥å¿µã«è¡ãªã£ãŠäžãã; ããã€ãã®ãšã³ã³ãŒãã£ã³ã°ã¯åé¡ãšãªãæ§è³ªãæã£ãŠããŸããäŸãã°å šåå°ã§ãªãã£ãããå®å šã« ASCII äºæã§ãªããªã©ãå ¥åããŒã¿ããšã³ã³ãŒãã£ã³ã°ãæå®ããŠããå Žåã§ãããããŠäžããããªããªãæ»æè ã¯å·§ã¿ã«æªæããããã¹ãããšã³ã³ãŒãããæååã®äžã«é ãããšãã§ããããã§ãã
ãã¡ã€ã«ãšã³ã³ãŒãã£ã³ã°ã®å€æÂ¶
The StreamRecoder class can transparently convert between
encodings, taking a stream that returns data in encoding #1
and behaving like a stream returning data in encoding #2.
For example, if you have an input file f that's in Latin-1, you
can wrap it with a StreamRecoder to return bytes encoded in
UTF-8:
new_f = codecs.StreamRecoder(f,
# en/decoder: used by read() to encode its results and
# by write() to decode its input.
codecs.getencoder('utf-8'), codecs.getdecoder('utf-8'),
# reader/writer: used to read and write to the stream.
codecs.getreader('latin-1'), codecs.getwriter('latin-1') )
äžæãªãšã³ã³ãŒãã£ã³ã°ã®ãã¡ã€ã«Â¶
What can you do if you need to make a change to a file, but don't know
the file's encoding? If you know the encoding is ASCII-compatible and
only want to examine or modify the ASCII parts, you can open the file
with the surrogateescape error handler:
with open(fname, 'r', encoding="ascii", errors="surrogateescape") as f:
data = f.read()
# make changes to the string 'data'
with open(fname + '.new', 'w',
encoding="ascii", errors="surrogateescape") as f:
f.write(data)
The surrogateescape error handler will decode any non-ASCII bytes
as code points in a special range running from U+DC80 to
U+DCFF. These code points will then turn back into the
same bytes when the surrogateescape error handler is used to
encode the data and write it back out.
åèè³æÂ¶
One section of Mastering Python 3 Input/Output, a PyCon 2010 talk by David Beazley, discusses text processing and binary data handling.
Marc-André Lemburg ã®ãã¬ãŒã³ããŒã·ã§ã³ "Writing Unicode-aware Applications in Python" ã® PDF ã¹ã©ã€ãã <https://downloads.egenix.com/python/LSM2005-Developing-Unicode-aware-applications-in-Python.pdf> ããå ¥æå¯èœã§ãããããŠæåãšã³ã³ãŒãã£ã³ã°ã®åé¡ãšåæ§ã«ã¢ããªã±ãŒã·ã§ã³ã®åœéåãããŒã«ã©ã€ãºã«ã€ããŠãè°è«ãããŠããŸãããã®ã¹ã©ã€ã㯠Python 2.x ã®ã¿ãã«ããŒããŠããŸãã
The Guts of Unicode in Python is a PyCon 2013 talk by Benjamin Peterson that discusses the internal Unicode representation in Python 3.3.
è¬èŸÂ¶
ãã®ããã¥ã¡ã³ãã®æåã®è皿㯠Andrew Kuchling ã«ãã£ãŠæžãããŸãããããããããã« Alexander Belopolsky, Georg Brandl, Andrew Kuchling, Ezio Melotti ãã§æ¹èšãéããããŠããŸãã
ãã®èšäºäžã®èª€ãã®ææãææ¡ãç³ãåºãŠããã以äžã®äººã ã«æè¬ããŸã: Ãric Araujo, Nicholas Bastin, Nick Coghlan, Marius Gedminas, Kent Johnson, Ken Krugler, Marc-André Lemburg, Martin von Löwis, Terry J. Reedy, Serhiy Storchaka, Eryk Sun, Chad Whitacre, Graham Wideman.