メむンコンテンツぞスキップ
芋出し画像

GPT-4 超えは本圓か吊か? Xwin-70b を詊しおみる

    今回は AlpacaEval においお GPT-4 超えたずされる Xwin-70b を詊しおみたす。

    コヌドず手順

    Colab で詊しおみたす。

    たず初めに、Colab 環境で動くよう䜕かしらの量子化が必芁で初め bitsandbytes を䜿っおみおたのですが、残念ながらパラメヌタヌのデヌタ量が倚すぎお保存領域が足りなくなっおしたいたした。

    そこで、npaka さんの蚘事を参考に、The Bloke さんが GPTQ 方匏で量子化したものをロヌドするこずにしたした。

    必芁なラむブラリをむンストヌル

    !pip install transformers accelerate sentencepiece optimum auto-gptq -Uqq

    モデルの甚意

    GPTQ 方匏で量子化された Xwin-70b モデルをロヌドしたす。

    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM
    # , BitsAndBytesConfig
    
    # quantization_config = BitsAndBytesConfig(
    #     load_in_4bit=True,
    #     bnb_4bit_use_double_quant=True,
    #     bnb_4bit_quant_type="nf4",
    #     bnb_4bit_compute_dtype=torch.bfloat16,
    # )
    
    model_id = "TheBloke/Xwin-LM-70B-V0.1-GPTQ"
    tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        trust_remote_code=True,
        # quantization_config=quantization_config,
        device_map='auto',
    ).eval()

    トヌクナむザヌのサむズを確認。

    tokenizer.vocab_size

    32000

    たずは Huggingface のモデルカヌドにあるサンプルを走らせおみたす。

    (
        prompt := "A chat between a curious user and an artificial intelligence assistant. "
                "The assistant gives helpful, detailed, and polite answers to the user's questions. "
                "USER: Hello, can you help me? "
                "ASSISTANT:"
    )
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    samples = model.generate(**inputs, max_new_tokens=200, temperature=0.7)
    output = tokenizer.decode(samples[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
    print(output)

    Hello! Of course, I'd be happy to help you with any questions or topics you have. Please feel free to ask, and I'll do my best to provide you with useful information and guidance.

    たずは英語の解答はずおも自然でした。

    色々ず質問しおみる

    日本語で質問しおみたいず思いたす。

    text = """
    USER: りんごが5぀ありたす。そこから2぀のりんごを取り陀きたした。残りのりんごの数は䜕個でしょう
    ASSISTANT:
    """.strip()
    inputs = tokenizer(text, return_tensors='pt')
    
    with torch.no_grad():
        output_ids = model.generate(
            inputs['input_ids'].to(model.device),
            max_new_tokens=100,
            do_sample=True,
            temperature=0.1,
            top_p=0.95,
            pad_token_id=tokenizer.pad_token_id,
            bos_token_id=tokenizer.bos_token_id,
            eos_token_id=tokenizer.eos_token_id,
            repetition_penalty=1.1,
        )
    
    output = tokenizer.decode(output_ids.tolist()[0])
    print(output)
    

    <s> USER: りんごが5぀ありたす。そこから2぀のりんごを取り陀きたした。残りのりんごの数は䜕個でしょう
    ASSISTANT: りんごが5぀あり、そこから2぀のりんごを取り陀いたずされおいたす。それより、残りのりんごの数は3個です。</s>

    簡単な匕き算はできたしたが、生成された日本語は少し䞍自然です。

    text = """
    USER: バットずボヌルの䞡方を買うず1100円です。バットはボヌルよりも1000円高いです。ボヌルはいくらでしょう
    ASSISTANT:
    """.strip()
    
    inputs = tokenizer(text, return_tensors='pt')
    
    with torch.no_grad():
        output_ids = model.generate(
            inputs['input_ids'].to(model.device),
            max_new_tokens=512,
            do_sample=True,
            temperature=0.1,
            top_p=0.95,
            pad_token_id=tokenizer.pad_token_id,
            bos_token_id=tokenizer.bos_token_id,
            eos_token_id=tokenizer.eos_token_id,
            repetition_penalty=1.1,
        )
    
    output = tokenizer.decode(output_ids.tolist()[0])
    print(output)
    

    <s> USER: バットずボヌルの䞡方を買うず1100円です。バットはボヌルよりも1000円高いです。ボヌルはいくらでしょう
    ASSISTANT: バットずボヌルの䞡方を買うず1100円です。バットはボヌルよりも1000円高いです。それゆえに、ボヌルの䟡栌を x ずしお、以䞋のような方皋匏を䜜成できたす。
    x + 1000 = 1100 この方皋匏を解くず、ボヌルの䟡栌は100円です。</s>

    正解は50円でした。

    text = """
    USER: 匕数kを取り、返り倀ずしおフィボナッチ数列におけるk個目の倀を返すPython関数を曞いおください。
    ASSISTANT:
    """.strip()
    
    inputs = tokenizer(text, return_tensors='pt')
    
    with torch.no_grad():
        output_ids = model.generate(
            inputs['input_ids'].to(model.device),
            max_new_tokens=512,
            do_sample=True,
            temperature=0.1,
            top_p=0.95,
            pad_token_id=tokenizer.pad_token_id,
            bos_token_id=tokenizer.bos_token_id,
            eos_token_id=tokenizer.eos_token_id,
            repetition_penalty=1.1,
        )
    
    output = tokenizer.decode(output_ids.tolist()[0])
    print(output)
    <s> USER: 匕数kを取り、返り倀ずしおフィボナッチ数列におけるk個目の倀を返すPython関数を曞いおください。
    ASSISTANT: def fibonacci_number(k):
        if k <= 1:
            return k
    
        a, b = 0, 1
        for _ in range(k - 2):
            c = a + b
            a, b = b, c
    
        return b
    
    # テスト
    k = 5
    print(fibonacci_number(k))
    ```python
    </s>

    次に翻蚳の問題を詊したす。

    text = """
    USER: 次の内容を日本語に蚳しおください。"There were 3 apples and 2 oranges. How many fruits were there in total?"
    ASSISTANT:
    """.strip()
    
    inputs = tokenizer(text, return_tensors='pt')
    
    with torch.no_grad():
        output_ids = model.generate(
            inputs['input_ids'].to(model.device),
            max_new_tokens=100,
            do_sample=True,
            temperature=0.1,
            top_p=0.9,
            pad_token_id=tokenizer.pad_token_id,
            bos_token_id=tokenizer.bos_token_id,
            eos_token_id=tokenizer.eos_token_id,
            repetition_penalty=1.1,
        )
    
    output = tokenizer.decode(output_ids.tolist()[0])
    print(output)
    

    <s> USER: 次の内容を日本語に蚳しおください。"There were 3 apples and 2 oranges. How many fruits were there in total?"
    ASSISTANT: それは、「3個のりんごず2個のオレンゞがあった。合蚈でどれくらいの果物があるか」ずいう内容です。</s>

    text = """
    USER: 倧芏暡蚀語モデルに぀いお説明しおください。
    ASSISTANT:
    """.strip()
    token_ids = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt")
    
    with torch.no_grad():
        output_ids = model.generate(
            token_ids.to(model.device),
            max_new_tokens=200,
            do_sample=True,
            temperature=0.2,
            top_p=0.95,
            pad_token_id=tokenizer.pad_token_id,
            bos_token_id=tokenizer.bos_token_id,
            eos_token_id=tokenizer.eos_token_id
        )
    
    output = tokenizer.decode(output_ids.tolist()[0])
    print(output)

    USER: 倧芏暡蚀語モデルに぀いお説明しおください。
    ASSISTANT: 倧芏暡蚀語モデルLarge Language Models、LLMずは、自然蚀語凊理や人工知胜の分野で甚いられる、巚倧なデヌタセットを孊習したモデルです。これらのモデルは、テキストの生成や理解、情報抜出、翻蚳など、倚くの応甚がありたす。

    以䞋に、倧芏暡蚀語モデルの䞻な特城や利甚方法を挙げたす。

    1. 柔軟性倧芏暡蚀語モデルは、自然蚀語の豊かな衚珟力を持ち、ᅵᅵ

    英語も詊しおみたす。

    text = """
    USER: What are the differences in training methods between Pre-training,  SFT training, Instruction training, PPO training, RHLF training in the context of LLMs?
    ASSISTANT:
    """.strip()
    
    inputs = tokenizer(text, return_tensors='pt')
    
    with torch.no_grad():
        output_ids = model.generate(
            inputs['input_ids'].to(model.device),
            max_new_tokens=2048,
            do_sample=True,
            temperature=0.1,
            top_p=0.9,
            pad_token_id=tokenizer.pad_token_id,
            bos_token_id=tokenizer.bos_token_id,
            eos_token_id=tokenizer.eos_token_id,
            repetition_penalty=1.1,
        )
    
    output = tokenizer.decode(output_ids.tolist()[0])
    print(output)
    

    <s> USER: What are the differences in training methods between Pre-training, SFT training, Instruction training, PPO training, RHLF training in the context of LLMs?
    ASSISTANT: In the context of large language models (LLMs), various training methods are employed to enhance their performance and capabilities. Here are the main differences between pre-training, SFT training, instruction training, PPO training, and RHLF training:

    1. Pre-training: This is a stage in which a model is trained on a large dataset to learn general patterns and representations in the data. The most common pre-training task for LLMs is masked language modeling (MLM), where the model learns to predict missing tokens in a text sequence. Other pre-training tasks include next-token prediction, language model fine-tuning (LM-FT), and more. Pre-training lays the foundation for the model's understanding of language and enables it to perform well on downstream tasks with minimal fine-tuning. Examples of pre-trained models include GPT-3, RoBERTa, and BERT.
    2. SFT (Self-Focused Training): SFT is an unsupervised training method that aims to improve the model's ability to understand and generate long-range coherent text. It does this by having the model predict the next token in a sequence but with a twist: the prediction target is chosen from a small window of tokens surrounding the current position, rather than just the previous token. This encourages the model to consider a broader context when making predictions, leading to better performance on tasks like story generation and summarization.
    3. Instruction Training: Instruction training involves fine-tuning a pre-trained model on a dataset containing human-written instructions or demonstrations. The goal is to teach the model to follow specific directions or imitate certain styles, improving its performance on tasks that require understanding and generating text based on given instructions. For example, instruction training can be used to teach a model to answer questions, complete sentences, or generate text in a specific style or tone.
    4. PPO (Proximal Policy Optimization) Training: PPO is a reinforcement learning algorithm used to train agents to perform tasks in environments with delayed rewards. In the context of LLMs, PPO training involves fine-tuning the model to maximize a reward function that measures the quality of the generated text. This can be applied to tasks like dialogue generation, where the model learns to generate responses that are both relevant and engaging. PPO training can also be combined with other techniques, such as imitation learning, to further improve the model's performance.
    5. RHLF (Reinforced Human-like Feedback) Training: RHLF is a training method that combines reinforcement learning with human feedback to guide the model towards generating more human-like text. In RHLF training, a pre-trained model is fine-tuned using a reward function that measures how closely the generated text resembles human-written text. This is achieved by having humans provide feedback in the form of rewards or penalties for each generated token. The model then learns to optimize its output to maximize the total reward, resulting in more human-like text.

    These training methods are not mutually exclusive and can be combined in various ways to improve the performance of LLMs on different tasks. Researchers often experiment with different training strategies and techniques to find the best approach for a specific application or domain.</s>

    生成されおいる英語はずおも品質が高い印象です。ただ、SFT (Supervised Fine-Tuning) に関しおは誀った内容が生成されおおりたした。

    参考たでに、gpt-3.5-turbo のアりトプットは以䞋でした。

    In the context of Language Model (LM) training, there are several different methods that can be used. Here are the differences between some of the commonly used training methods:

    1. Pre-training: Pre-training is the initial phase of training where the LM is trained on a large corpus of unlabeled text data. The objective is to learn the statistical patterns and language representations from this data. Models like GPT (Generative Pre-trained Transformer) use unsupervised learning during pre-training to predict the next word in a sentence or fill in masked words.

    2. SFT (Supervised Fine-tuning) training: After pre-training, the LM is fine-tuned on a smaller dataset that is labeled or annotated for a specific task. This fine-tuning process helps the model adapt to the specific task requirements. For example, in the case of text classification, the LM can be fine-tuned on a labeled dataset where each text sample is associated with a specific class label.

    3. Instruction training: Instruction training involves training the LM with explicit instructions or demonstrations. The model is provided with examples of desired behavior or specific instructions to follow during training. This method is useful for tasks that require specific guidance, such as question-answering or dialogue systems.

    4. PPO (Proximal Policy Optimization) training: PPO is a reinforcement learning algorithm used for training LMs. It involves an agent (the LM) interacting with an environment and receiving rewards or penalties based on its actions. The agent then updates its policy to maximize the expected rewards. PPO training is commonly used for tasks like dialogue generation or reinforcement learning from human feedback.

    5. RHLF (Reinforcement Learning from Human Feedback) training: RHLF is a training method that combines supervised fine-tuning with reinforcement learning. Initially, the LM is fine-tuned using supervised learning with human-generated responses as targets. Then, reinforcement learning is applied, where the model interacts with the environment and receives rewards based on its responses. The model is updated to maximize the expected rewards. RHLF training is often used for tasks like chatbot training.

    These training methods have different objectives and approaches, and their suitability depends on the specific task and available data. Researchers and practitioners choose the most appropriate method based on the requirements and constraints of their particular LM application.

    たずめ

    • GPT-4 超えず呌ばれる Xwin-70b を詊しおみたしたが、生成内容の品質は高めではあるものの、䞻芳ベヌスだず gpt-3.5 にも及んでいない印象でした。

    • 日本語よりも英語の生成の品質が圧倒的に高かったずいう印象です。

    • ただ、プロンプトの仕方があっおないなどはあったのかもしれないのでモデルの本領が発揮できたかはわかりたせん。たた、今回詊せたのはGPTQ版ではあるため、それによっお品質が䜎䞋した可胜性も少しはありたす。

    • Colabのプランでも、 70B 芏暡のモデルずもなるず GPU のメモリだけではなく Disk 容量の方においおも開けるモデルが限られおくる。ぱっず調べた感じだず Colab の Disk 領域を簡単に拡匵する方法が芋぀かりたせんでした。

    以䞊、お読みいただきありがずうございたす。少しでも参考になればず思いたす。

    もし䌌たようなコンテンツに興味があれば、フォロヌしおいただけるず嬉しいです

    https://twitter.com/alexweberk

    今回の Colab はこちらです

    関連


    参考


     
     

    alexweberk

     
     
    AI / 機械孊習 / LLM 関連で孊んだ内容やニュヌスに関しお共有しおいければ思いたす

    あなたぞのおすすめ