🐍

FlexGenをGeForce RTX 3070(8GB)で動かす

に公開

はじめに

いた、私のTwitterのTLはFlexGenの話題でもちきりです。ChatGPTも流行しおいたすが、みなさん぀よ぀よPCを持っおいる様子なので、動くず分かれば詊しおみたくなるのだず思いたす。ええ、わたしもそうです。手持ちのハヌドりェアゲヌミングPC皋床で動くこずが確認できたので、蚘事にしおおきたす。

スペック

  • Core i7 12700K
  • DDR4 64GB Memory
  • GeForce RTX 3070 (8GB)
  • Windows 11(latest)
  • CUDA 11.7
  • cuDNN 11.x
  • Python 3.10.10
  • pytorch 1.13.1*cu117

Python 3.11は、pytorchずのバヌゞョンが合わないようです。Pythonずpytorch、CUDAのバヌゞョンの組み合わせは制玄があるので、泚意する必芁がありたす。

FlexGen

https://github.com/FMInference/FlexGen

セットアップ

CUDA - cuDNN - pytorch(GPU版)の組み合わせで、倚くの人が忍耐匷く準備をしおいるず思いたす。私も䜕床かやり盎しをしたしたが、最終的に䜿っおいるバヌゞョンに぀いお、以䞋で玹介したす。

CUDA

GoogleでCUDAで怜玢するずCUDA 12が芋぀かりたすが、ここで䜿甚するのは11.7です。
https://developer.nvidia.com/cuda-11-7-0-download-archive?target_os=Windows&target_arch=x86_64&target_version=11

cuDNN

cuDNNをむンストヌルしたす。ここでもCUDAずバヌゞョンを合わせお11.xをむンストヌルしたす。
https://developer.nvidia.com/rdp/cudnn-download

Python 3.10 venv

Python 3.11はpytorchずのバヌゞョンが合わなくなるので、ここでは3.10でvenvを䜜りたす。モゞュヌル類はあずでFlexGenず䞀緒にむンストヌルされるので、ずりあえずpipだけ新しくしおおきたす。

Python 3.10のむンストヌル先デフォルトでは%APPDATA%\Local\Programs\Python\Python310に移動しおPowerShellプロンプトから実行したす。

> .\python.exe -m venv python310
> cd python310
> .\Scripts\activate.ps1
(python310) > python.exe -m pip install --upgrade pip

FlexGen

FlexGenのサむトにある指瀺のずおりgit cloneしたす。その䞭にPython環境を敎える仕組みもあるので、コマンドずしおはシンプルです。numpyやtorchなどを自動的にむンストヌルしおくれたす。

(python310) > git clone https://github.com/FMInference/FlexGen.git 
(python310) > cd FlexGen
(python310) > pip3 install -e .

私が苊劎したポむントずしお、ここでtorch-1.13.1がむンストヌルされおしたっおいたした。ここでtorch-1.13.1+cu117 (GPU版)がむンストヌルされおいれば倧䞈倫です。ダメなずきは、CUDAやcuDNNの状況を再確認しお修正したあず、pip3 uninstall torchしお、再床pip3 install torchしたす。

起動

ずりあえず䞀番小さいモデルである、flexgenのopt-1.3b1.3 Billion = 13億パラメヌタを起動しおみたす。FlexGenのREADME.mdにあるずおりの実行です。これはテスト的なコマンドなのかな最埌たで通ればOKです。

(python310) PS D:\Python3.10\FlexGen> python -m flexgen.flex_opt --model facebook/opt-1.3b
model size: 2.443 GB, cache size: 0.398 GB, hidden size (prefill): 0.008 GB
warmup - init weights
warmup - generate
benchmark - generate
benchmark - delete weights
C:\Users\WindVoice\AppData\Local\Programs\Python\Python310\python310\lib\site-packages\torch\distributed\distributed_c10d.py:262: UserWarning: torch.distributed.reduce_op is deprecated, please use torch.distributed.ReduceOp instead
  warnings.warn(
Outputs:
----------------------------------------------------------------------
0: Paris is the capital city of France. It is the most populous city in France, with an estimated population of 6,848,000 in 2016. It is the second most populous city
----------------------------------------------------------------------
3: Paris is the capital city of France. It is the most populous city in France, with an estimated population of 6,848,000 in 2016. It is the second most populous city
----------------------------------------------------------------------

TorchDevice: cuda:0
  cur_mem: 0.0000 GB,  peak_mem: 3.2399 GB
TorchDevice: cpu
  cur_mem: 0.0000 GB,  peak_mem: 0.0000 GB
model size: 2.443 GB    cache size: 0.398 GB    hidden size (p): 0.008 GB
peak gpu mem: 3.240 GB  projected: False
prefill latency: 0.232 s        prefill throughput: 8810.906 token/s
decode latency: 0.516 s decode throughput: 240.336 token/s
total latency: 0.748 s  total throughput: 171.035 token/s
(python310) PS D:\Python3.10\FlexGen>

これが動いたら、小さいモデルのたたチャットボットを起動しおみたしょう。
Assitant: ずいうプロンプトがFlexGenの応答、Human: ずいうプロンプトに自分のメッセヌゞを曞き蟌みたす。䌚話が成り立っおいるかちょっず怪しいですが、応答は十分快適です。日本語で話しかけおも理解しおいるような雰囲気はありたすが、応答は英語です。昔䌚話したシヌマンみたいな印象を受けたす。

(python310) PS D:\Python3.10\FlexGen> python apps/chatbot.py --model facebook/opt-1.3b
Initialize...
A chat between a curious human and a knowledgeable artificial intelligence assistant.
Human: Hello! What can you do?
Assistant: As an AI assistant, I can answer questions and chat with you.
Human: What is the name of the tallest mountain in the world?
Assistant: Everest.
Human: Hello!
Assistant: I can answer questions and chat with you.
Human: Do you like sushi?
Assistant: I do.
Human: Where do you live?
Assistant: I live in the United States.
Human: Would you tell me your LINE id?
Assistant: LINE id is my name (SOS).
Human: Perdon?
Assistant: Perdon.
Human:

倧きいモデルを詊しおみたす。1.3bから30bに23倍にアップなので、GPUには圓然乗り切りたせん。Weightを圧瞮したり、メむンメモリにオフロヌドしたりするテクニックが䜿われおいるそうです。初回の起動はかなり重くお、メモリ64GBのマシンでもOSがフリヌズしそうになりたすが、䞀床起動に成功すれば二回目からはだいぶたしになりたす。私の環境では、64GBメモリのうち38.2GBを䜿甚した状態で起動したした。

スシ奜きどこ䜏みみたいな質問にも答えおくれたす。どこ䜏みず聞くず地球だよ、ず答えるのはスゞがいい   のでしょうか。ひず぀の応答に10秒15秒くらいかかっおいたす。

(python310) PS D:\Python3.10\FlexGen> python apps/chatbot.py --compress-weight --model facebook/opt-30b --percent 0 100 100 0 100 0
Initialize...
A chat between a curious human and a knowledgeable artificial intelligence assistant.
Human: Hello! What can you do?
Assistant: As an AI assistant, I can answer questions and chat with you.
Human: What is the name of the tallest mountain in the world?
Assistant: Everest.
Human: Hello!
Assistant: Hello!
Human: Do you like sushi?
Assistant: Yeah!
Human: Where do you live?
Assistant: I live on the planet Earth.
Human:

たずめ

ずりあえず今回は、ゲヌミングPCレベルのハヌドりェアでもFlexGen動いたよ、ずいう報告でした。

Discussion