FlexGenãGeForce RTX 3070(8GB)ã§åãã
ã¯ããã«
ããŸãç§ã®Twitterã®TLã¯FlexGenã®è©±é¡ã§ãã¡ããã§ããChatGPTãæµè¡ããŠããŸãããã¿ãªããã€ãã€ãPCãæã£ãŠããæ§åãªã®ã§ãåããšåããã°è©ŠããŠã¿ãããªãã®ã ãšæããŸãããããããããããã§ããææã¡ã®ããŒããŠã§ã¢ïŒã²ãŒãã³ã°PCçšåºŠïŒã§åãããšã確èªã§ããã®ã§ãèšäºã«ããŠãããŸãã
ã¹ããã¯
- Core i7 12700K
- DDR4 64GB Memory
- GeForce RTX 3070 (8GB)
- Windows 11(latest)
- CUDA 11.7
- cuDNN 11.x
- Python 3.10.10
- pytorch 1.13.1*cu117
Python 3.11ã¯ãpytorchãšã®ããŒãžã§ã³ãåããªãããã§ããPythonãšpytorchãCUDAã®ããŒãžã§ã³ã®çµã¿åããã¯å¶çŽãããã®ã§ã泚æããå¿ èŠããããŸãã
FlexGen
ã»ããã¢ãã
CUDA - cuDNN - pytorch(GPUç)ã®çµã¿åããã§ãå€ãã®äººãå¿èåŒ·ãæºåãããŠãããšæããŸããç§ãäœåºŠãããçŽããããŸããããæçµçã«äœ¿ã£ãŠããããŒãžã§ã³ã«ã€ããŠã以äžã§ç޹ä»ããŸãã
CUDA
Googleã§CUDAã§æ€çŽ¢ãããšCUDA 12ãèŠã€ãããŸãããããã§äœ¿çšããã®ã¯11.7ã§ãã
cuDNN
cuDNNãã€ã³ã¹ããŒã«ããŸããããã§ãCUDAãšããŒãžã§ã³ãåãããŠ11.xãã€ã³ã¹ããŒã«ããŸãã
Python 3.10 venv
Python 3.11ã¯pytorchãšã®ããŒãžã§ã³ãåããªããªãã®ã§ãããã§ã¯3.10ã§venvãäœããŸããã¢ãžã¥ãŒã«é¡ã¯ããšã§FlexGenãšäžç·ã«ã€ã³ã¹ããŒã«ãããã®ã§ããšããããpipã ãæ°ããããŠãããŸãã
Python 3.10ã®ã€ã³ã¹ããŒã«å ïŒããã©ã«ãã§ã¯%APPDATA%\Local\Programs\Python\Python310ïŒã«ç§»åããŠPowerShellããã³ããããå®è¡ããŸãã
> .\python.exe -m venv python310
> cd python310
> .\Scripts\activate.ps1
(python310) > python.exe -m pip install --upgrade pip
FlexGen
FlexGenã®ãµã€ãã«ããæç€ºã®ãšããgit cloneããŸãããã®äžã«Pythonç°å¢ãæŽããä»çµã¿ãããã®ã§ãã³ãã³ããšããŠã¯ã·ã³ãã«ã§ããnumpyãtorchãªã©ãèªåçã«ã€ã³ã¹ããŒã«ããŠãããŸãã
(python310) > git clone https://github.com/FMInference/FlexGen.git
(python310) > cd FlexGen
(python310) > pip3 install -e .
ç§ãèŠåŽãããã€ã³ããšããŠãããã§torch-1.13.1ãã€ã³ã¹ããŒã«ãããŠããŸã£ãŠããŸãããããã§torch-1.13.1+cu117 (GPUç)ãã€ã³ã¹ããŒã«ãããŠããã°å€§äžå€«ã§ãããã¡ãªãšãã¯ãCUDAãcuDNNã®ç¶æ³ãå確èªããŠä¿®æ£ããããšãpip3 uninstall torchããŠãå床pip3 install torchããŸãã
èµ·åïŒ
ãšããããäžçªå°ããã¢ãã«ã§ãããflexgenã®opt-1.3bïŒ1.3 Billion = 13åãã©ã¡ãŒã¿ïŒãèµ·åããŠã¿ãŸããFlexGenã®README.mdã«ãããšããã®å®è¡ã§ããããã¯ãã¹ãçãªã³ãã³ããªã®ããªïŒæåŸãŸã§éãã°OKã§ãã
(python310) PS D:\Python3.10\FlexGen> python -m flexgen.flex_opt --model facebook/opt-1.3b
model size: 2.443 GB, cache size: 0.398 GB, hidden size (prefill): 0.008 GB
warmup - init weights
warmup - generate
benchmark - generate
benchmark - delete weights
C:\Users\WindVoice\AppData\Local\Programs\Python\Python310\python310\lib\site-packages\torch\distributed\distributed_c10d.py:262: UserWarning: torch.distributed.reduce_op is deprecated, please use torch.distributed.ReduceOp instead
warnings.warn(
Outputs:
----------------------------------------------------------------------
0: Paris is the capital city of France. It is the most populous city in France, with an estimated population of 6,848,000 in 2016. It is the second most populous city
----------------------------------------------------------------------
3: Paris is the capital city of France. It is the most populous city in France, with an estimated population of 6,848,000 in 2016. It is the second most populous city
----------------------------------------------------------------------
TorchDevice: cuda:0
cur_mem: 0.0000 GB, peak_mem: 3.2399 GB
TorchDevice: cpu
cur_mem: 0.0000 GB, peak_mem: 0.0000 GB
model size: 2.443 GB cache size: 0.398 GB hidden size (p): 0.008 GB
peak gpu mem: 3.240 GB projected: False
prefill latency: 0.232 s prefill throughput: 8810.906 token/s
decode latency: 0.516 s decode throughput: 240.336 token/s
total latency: 0.748 s total throughput: 171.035 token/s
(python310) PS D:\Python3.10\FlexGen>
ãããåããããå°ããã¢ãã«ã®ãŸãŸãã£ããããããèµ·åããŠã¿ãŸãããã
Assitant: ãšããããã³ãããFlexGenã®å¿çãHuman: ãšããããã³ããã«èªåã®ã¡ãã»ãŒãžãæžã蟌ã¿ãŸããäŒè©±ãæãç«ã£ãŠãããã¡ãã£ãšæªããã§ãããå¿çã¯ååå¿«é©ã§ããæ¥æ¬èªã§è©±ããããŠãçè§£ããŠãããããªé°å²æ°ã¯ãããŸãããå¿çã¯è±èªã§ããæäŒè©±ããã·ãŒãã³ã¿ãããªå°è±¡ãåããŸãã
(python310) PS D:\Python3.10\FlexGen> python apps/chatbot.py --model facebook/opt-1.3b
Initialize...
A chat between a curious human and a knowledgeable artificial intelligence assistant.
Human: Hello! What can you do?
Assistant: As an AI assistant, I can answer questions and chat with you.
Human: What is the name of the tallest mountain in the world?
Assistant: Everest.
Human: Hello!
Assistant: I can answer questions and chat with you.
Human: Do you like sushi?
Assistant: I do.
Human: Where do you live?
Assistant: I live in the United States.
Human: Would you tell me your LINE id?
Assistant: LINE id is my name (SOS).
Human: Perdon?
Assistant: Perdon.
Human:
倧ããã¢ãã«ã詊ããŠã¿ãŸãã1.3bãã30bã«23åã«ã¢ãããªã®ã§ãGPUã«ã¯åœç¶ä¹ãåããŸãããWeightãå§çž®ããããã¡ã€ã³ã¡ã¢ãªã«ãªãããŒãããããããã¯ããã¯ã䜿ãããŠããããã§ããååã®èµ·åã¯ããªãéããŠãã¡ã¢ãª64GBã®ãã·ã³ã§ãOSãããªãŒãºãããã«ãªããŸãããäžåºŠèµ·åã«æåããã°äºåç®ããã¯ã ãã¶ãŸãã«ãªããŸããç§ã®ç°å¢ã§ã¯ã64GBã¡ã¢ãªã®ãã¡38.2GBã䜿çšããç¶æ ã§èµ·åããŸããã
ã¹ã·å¥œãïŒã©ãäœã¿ïŒã¿ãããªè³ªåã«ãçããŠãããŸããã©ãäœã¿ïŒãšèããšå°çã ãããšçããã®ã¯ã¹ãžãããâŠâŠãã®ã§ãããããã²ãšã€ã®å¿çã«10ç§ïœ15ç§ãããããã£ãŠããŸãã
(python310) PS D:\Python3.10\FlexGen> python apps/chatbot.py --compress-weight --model facebook/opt-30b --percent 0 100 100 0 100 0
Initialize...
A chat between a curious human and a knowledgeable artificial intelligence assistant.
Human: Hello! What can you do?
Assistant: As an AI assistant, I can answer questions and chat with you.
Human: What is the name of the tallest mountain in the world?
Assistant: Everest.
Human: Hello!
Assistant: Hello!
Human: Do you like sushi?
Assistant: Yeah!
Human: Where do you live?
Assistant: I live on the planet Earth.
Human:
ãŸãšã
ãšããããä»åã¯ãã²ãŒãã³ã°PCã¬ãã«ã®ããŒããŠã§ã¢ã§ãFlexGenåãããããšããå ±åã§ããã
Discussion