
GPT-4 è¶ ãã¯æ¬åœãåŠã? Xwin-70b ã詊ããŠã¿ã
ä»å㯠AlpacaEval ã«ãã㊠GPT-4 è¶ ãããšããã Xwin-70b ã詊ããŠã¿ãŸãã
Huggingface: https://huggingface.co/Xwin-LM/Xwin-LM-70B-V0.1
è«æ: æªçºè¡š
ã©ã€ã»ã³ã¹: Llama 2 License
ã³ãŒããšæé
Colab ã§è©ŠããŠã¿ãŸãã
ãŸãåãã«ãColab ç°å¢ã§åãããäœãããã®éååãå¿ èŠã§åã bitsandbytes ã䜿ã£ãŠã¿ãŠãã®ã§ãããæ®å¿µãªãããã©ã¡ãŒã¿ãŒã®ããŒã¿éãå€ãããŠä¿åé åãè¶³ããªããªã£ãŠããŸããŸããã
ããã§ãnpaka ããã®èšäºãåèã«ãThe Bloke ããã GPTQ æ¹åŒã§éååãããã®ãããŒãããããšã«ããŸããã
å¿ èŠãªã©ã€ãã©ãªãã€ã³ã¹ããŒã«
!pip install transformers accelerate sentencepiece optimum auto-gptq -Uqqã¢ãã«ã®çšæ
GPTQ æ¹åŒã§éååããã Xwin-70b ã¢ãã«ãããŒãããŸãã
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# , BitsAndBytesConfig
# quantization_config = BitsAndBytesConfig(
# load_in_4bit=True,
# bnb_4bit_use_double_quant=True,
# bnb_4bit_quant_type="nf4",
# bnb_4bit_compute_dtype=torch.bfloat16,
# )
model_id = "TheBloke/Xwin-LM-70B-V0.1-GPTQ"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
# quantization_config=quantization_config,
device_map='auto',
).eval()ããŒã¯ãã€ã¶ãŒã®ãµã€ãºã確èªã
tokenizer.vocab_size32000
ãŸã㯠Huggingface ã®ã¢ãã«ã«ãŒãã«ãããµã³ãã«ãèµ°ãããŠã¿ãŸãã
(
prompt := "A chat between a curious user and an artificial intelligence assistant. "
"The assistant gives helpful, detailed, and polite answers to the user's questions. "
"USER: Hello, can you help me? "
"ASSISTANT:"
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
samples = model.generate(**inputs, max_new_tokens=200, temperature=0.7)
output = tokenizer.decode(samples[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(output)Hello! Of course, I'd be happy to help you with any questions or topics you have. Please feel free to ask, and I'll do my best to provide you with useful information and guidance.
ãŸãã¯è±èªã®è§£çã¯ãšãŠãèªç¶ã§ããã
è²ã ãšè³ªåããŠã¿ã
æ¥æ¬èªã§è³ªåããŠã¿ãããšæããŸãã
text = """
USER: ãããã5ã€ãããŸãããããã2ã€ã®ããããåãé€ããŸãããæ®ãã®ãããã®æ°ã¯äœåã§ãããïŒ
ASSISTANT:
""".strip()
inputs = tokenizer(text, return_tensors='pt')
with torch.no_grad():
output_ids = model.generate(
inputs['input_ids'].to(model.device),
max_new_tokens=100,
do_sample=True,
temperature=0.1,
top_p=0.95,
pad_token_id=tokenizer.pad_token_id,
bos_token_id=tokenizer.bos_token_id,
eos_token_id=tokenizer.eos_token_id,
repetition_penalty=1.1,
)
output = tokenizer.decode(output_ids.tolist()[0])
print(output)
<s> USER: ãããã5ã€ãããŸãããããã2ã€ã®ããããåãé€ããŸãããæ®ãã®ãããã®æ°ã¯äœåã§ãããïŒ
ASSISTANT: ãããã5ã€ããããããã2ã€ã®ããããåãé€ãããšãããŠããŸãããããããæ®ãã®ãããã®æ°ã¯3åã§ãã</s>
ç°¡åãªåŒãç®ã¯ã§ããŸããããçæãããæ¥æ¬èªã¯å°ãäžèªç¶ã§ãã
text = """
USER: ããããšããŒã«ã®äž¡æ¹ãè²·ããš1100åã§ãããããã¯ããŒã«ããã1000åé«ãã§ããããŒã«ã¯ãããã§ãããïŒ
ASSISTANT:
""".strip()
inputs = tokenizer(text, return_tensors='pt')
with torch.no_grad():
output_ids = model.generate(
inputs['input_ids'].to(model.device),
max_new_tokens=512,
do_sample=True,
temperature=0.1,
top_p=0.95,
pad_token_id=tokenizer.pad_token_id,
bos_token_id=tokenizer.bos_token_id,
eos_token_id=tokenizer.eos_token_id,
repetition_penalty=1.1,
)
output = tokenizer.decode(output_ids.tolist()[0])
print(output)
<s> USER: ããããšããŒã«ã®äž¡æ¹ãè²·ããš1100åã§ãããããã¯ããŒã«ããã1000åé«ãã§ããããŒã«ã¯ãããã§ãããïŒ
ASSISTANT: ããããšããŒã«ã®äž¡æ¹ãè²·ããš1100åã§ãããããã¯ããŒã«ããã1000åé«ãã§ããããããã«ãããŒã«ã®äŸ¡æ Œã x ãšããŠã以äžã®ãããªæ¹çšåŒãäœæã§ããŸãã
x + 1000 = 1100 ãã®æ¹çšåŒãè§£ããšãããŒã«ã®äŸ¡æ Œã¯100åã§ãã</s>
æ£è§£ã¯50åã§ããã
text = """
USER: åŒæ°kãåããè¿ãå€ãšããŠãã£ããããæ°åã«ãããkåç®ã®å€ãè¿ãPython颿°ãæžããŠãã ããã
ASSISTANT:
""".strip()
inputs = tokenizer(text, return_tensors='pt')
with torch.no_grad():
output_ids = model.generate(
inputs['input_ids'].to(model.device),
max_new_tokens=512,
do_sample=True,
temperature=0.1,
top_p=0.95,
pad_token_id=tokenizer.pad_token_id,
bos_token_id=tokenizer.bos_token_id,
eos_token_id=tokenizer.eos_token_id,
repetition_penalty=1.1,
)
output = tokenizer.decode(output_ids.tolist()[0])
print(output)<s> USER: åŒæ°kãåããè¿ãå€ãšããŠãã£ããããæ°åã«ãããkåç®ã®å€ãè¿ãPython颿°ãæžããŠãã ããã
ASSISTANT: def fibonacci_number(k):
if k <= 1:
return k
a, b = 0, 1
for _ in range(k - 2):
c = a + b
a, b = b, c
return b
# ãã¹ã
k = 5
print(fibonacci_number(k))
```python
</s>次ã«ç¿»èš³ã®åé¡ã詊ããŸãã
text = """
USER: 次ã®å
å®¹ãæ¥æ¬èªã«èš³ããŠãã ããã"There were 3 apples and 2 oranges. How many fruits were there in total?"
ASSISTANT:
""".strip()
inputs = tokenizer(text, return_tensors='pt')
with torch.no_grad():
output_ids = model.generate(
inputs['input_ids'].to(model.device),
max_new_tokens=100,
do_sample=True,
temperature=0.1,
top_p=0.9,
pad_token_id=tokenizer.pad_token_id,
bos_token_id=tokenizer.bos_token_id,
eos_token_id=tokenizer.eos_token_id,
repetition_penalty=1.1,
)
output = tokenizer.decode(output_ids.tolist()[0])
print(output)
<s> USER: 次ã®å å®¹ãæ¥æ¬èªã«èš³ããŠãã ããã"There were 3 apples and 2 oranges. How many fruits were there in total?"
ASSISTANT: ããã¯ãã3åã®ããããš2åã®ãªã¬ã³ãžããã£ããåèšã§ã©ããããã®æç©ããããïŒããšããå 容ã§ãã</s>
text = """
USER: å€§èŠæš¡èšèªã¢ãã«ã«ã€ããŠèª¬æããŠãã ããã
ASSISTANT:
""".strip()
token_ids = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt")
with torch.no_grad():
output_ids = model.generate(
token_ids.to(model.device),
max_new_tokens=200,
do_sample=True,
temperature=0.2,
top_p=0.95,
pad_token_id=tokenizer.pad_token_id,
bos_token_id=tokenizer.bos_token_id,
eos_token_id=tokenizer.eos_token_id
)
output = tokenizer.decode(output_ids.tolist()[0])
print(output)USER: å€§èŠæš¡èšèªã¢ãã«ã«ã€ããŠèª¬æããŠãã ããã
ASSISTANT: å€§èŠæš¡èšèªã¢ãã«ïŒLarge Language ModelsãLLMïŒãšã¯ãèªç¶èšèªåŠçã人工ç¥èœã®åéã§çšããããã巚倧ãªããŒã¿ã»ãããåŠç¿ããã¢ãã«ã§ãããããã®ã¢ãã«ã¯ãããã¹ãã®çæãçè§£ãæ å ±æœåºã翻蚳ãªã©ãå€ãã®å¿çšããããŸãã
以äžã«ãå€§èŠæš¡èšèªã¢ãã«ã®äž»ãªç¹åŸŽãå©ç𿹿³ãæããŸãã
1. æè»æ§ïŒå€§èŠæš¡èšèªã¢ãã«ã¯ãèªç¶èšèªã®è±ããªè¡šçŸåãæã¡ãᅵᅵ
è±èªã詊ããŠã¿ãŸãã
text = """
USER: What are the differences in training methods between Pre-training, SFT training, Instruction training, PPO training, RHLF training in the context of LLMs?
ASSISTANT:
""".strip()
inputs = tokenizer(text, return_tensors='pt')
with torch.no_grad():
output_ids = model.generate(
inputs['input_ids'].to(model.device),
max_new_tokens=2048,
do_sample=True,
temperature=0.1,
top_p=0.9,
pad_token_id=tokenizer.pad_token_id,
bos_token_id=tokenizer.bos_token_id,
eos_token_id=tokenizer.eos_token_id,
repetition_penalty=1.1,
)
output = tokenizer.decode(output_ids.tolist()[0])
print(output)
<s> USER: What are the differences in training methods between Pre-training, SFT training, Instruction training, PPO training, RHLF training in the context of LLMs?
ASSISTANT: In the context of large language models (LLMs), various training methods are employed to enhance their performance and capabilities. Here are the main differences between pre-training, SFT training, instruction training, PPO training, and RHLF training:
1. Pre-training: This is a stage in which a model is trained on a large dataset to learn general patterns and representations in the data. The most common pre-training task for LLMs is masked language modeling (MLM), where the model learns to predict missing tokens in a text sequence. Other pre-training tasks include next-token prediction, language model fine-tuning (LM-FT), and more. Pre-training lays the foundation for the model's understanding of language and enables it to perform well on downstream tasks with minimal fine-tuning. Examples of pre-trained models include GPT-3, RoBERTa, and BERT.
2. SFT (Self-Focused Training): SFT is an unsupervised training method that aims to improve the model's ability to understand and generate long-range coherent text. It does this by having the model predict the next token in a sequence but with a twist: the prediction target is chosen from a small window of tokens surrounding the current position, rather than just the previous token. This encourages the model to consider a broader context when making predictions, leading to better performance on tasks like story generation and summarization.
3. Instruction Training: Instruction training involves fine-tuning a pre-trained model on a dataset containing human-written instructions or demonstrations. The goal is to teach the model to follow specific directions or imitate certain styles, improving its performance on tasks that require understanding and generating text based on given instructions. For example, instruction training can be used to teach a model to answer questions, complete sentences, or generate text in a specific style or tone.
4. PPO (Proximal Policy Optimization) Training: PPO is a reinforcement learning algorithm used to train agents to perform tasks in environments with delayed rewards. In the context of LLMs, PPO training involves fine-tuning the model to maximize a reward function that measures the quality of the generated text. This can be applied to tasks like dialogue generation, where the model learns to generate responses that are both relevant and engaging. PPO training can also be combined with other techniques, such as imitation learning, to further improve the model's performance.
5. RHLF (Reinforced Human-like Feedback) Training: RHLF is a training method that combines reinforcement learning with human feedback to guide the model towards generating more human-like text. In RHLF training, a pre-trained model is fine-tuned using a reward function that measures how closely the generated text resembles human-written text. This is achieved by having humans provide feedback in the form of rewards or penalties for each generated token. The model then learns to optimize its output to maximize the total reward, resulting in more human-like text.
These training methods are not mutually exclusive and can be combined in various ways to improve the performance of LLMs on different tasks. Researchers often experiment with different training strategies and techniques to find the best approach for a specific application or domain.</s>
çæãããŠããè±èªã¯ãšãŠãå質ãé«ãå°è±¡ã§ãããã ãSFT (Supervised Fine-Tuning) ã«é¢ããŠã¯èª€ã£ãå 容ãçæãããŠãããŸããã
åèãŸã§ã«ãgpt-3.5-turbo ã®ã¢ãŠããããã¯ä»¥äžã§ããã
In the context of Language Model (LM) training, there are several different methods that can be used. Here are the differences between some of the commonly used training methods:
1. Pre-training: Pre-training is the initial phase of training where the LM is trained on a large corpus of unlabeled text data. The objective is to learn the statistical patterns and language representations from this data. Models like GPT (Generative Pre-trained Transformer) use unsupervised learning during pre-training to predict the next word in a sentence or fill in masked words.
2. SFT (Supervised Fine-tuning) training: After pre-training, the LM is fine-tuned on a smaller dataset that is labeled or annotated for a specific task. This fine-tuning process helps the model adapt to the specific task requirements. For example, in the case of text classification, the LM can be fine-tuned on a labeled dataset where each text sample is associated with a specific class label.
3. Instruction training: Instruction training involves training the LM with explicit instructions or demonstrations. The model is provided with examples of desired behavior or specific instructions to follow during training. This method is useful for tasks that require specific guidance, such as question-answering or dialogue systems.
4. PPO (Proximal Policy Optimization) training: PPO is a reinforcement learning algorithm used for training LMs. It involves an agent (the LM) interacting with an environment and receiving rewards or penalties based on its actions. The agent then updates its policy to maximize the expected rewards. PPO training is commonly used for tasks like dialogue generation or reinforcement learning from human feedback.
5. RHLF (Reinforcement Learning from Human Feedback) training: RHLF is a training method that combines supervised fine-tuning with reinforcement learning. Initially, the LM is fine-tuned using supervised learning with human-generated responses as targets. Then, reinforcement learning is applied, where the model interacts with the environment and receives rewards based on its responses. The model is updated to maximize the expected rewards. RHLF training is often used for tasks like chatbot training.
These training methods have different objectives and approaches, and their suitability depends on the specific task and available data. Researchers and practitioners choose the most appropriate method based on the requirements and constraints of their particular LM application.
ãŸãšã
GPT-4 è¶ ããšåŒã°ãã Xwin-70b ã詊ããŠã¿ãŸããããçæå 容ã®å質ã¯é«ãã§ã¯ãããã®ã®ã䞻芳ããŒã¹ã ãš gpt-3.5 ã«ãåãã§ããªãå°è±¡ã§ããã
æ¥æ¬èªãããè±èªã®çæã®å質ãå§åçã«é«ãã£ããšããå°è±¡ã§ãã
ãã ãããã³ããã®ä»æ¹ããã£ãŠãªããªã©ã¯ãã£ãã®ãããããªãã®ã§ã¢ãã«ã®æ¬é ãçºæ®ã§ãããã¯ããããŸããããŸããä»å詊ããã®ã¯GPTQçã§ã¯ãããããããã«ãã£ãŠå質ãäœäžããå¯èœæ§ãïŒå°ãã¯ïŒãããŸãã
ColabïŒã®ãã©ã³ã§ãã 70B èŠæš¡ã®ã¢ãã«ãšããªããš GPU ã®ã¡ã¢ãªã ãã§ã¯ãªã Disk 容éã®æ¹ã«ãããŠãéããã¢ãã«ãéãããŠãããã±ã£ãšèª¿ã¹ãæãã ãš Colab ã® Disk é åãç°¡åã«æ¡åŒµããæ¹æ³ãèŠã€ãããŸããã§ããã
以äžããèªã¿ããã ãããããšãããããŸããå°ãã§ãåèã«ãªãã°ãšæããŸãã
ãã䌌ããããªã³ã³ãã³ãã«èå³ãããã°ããã©ããŒããŠããã ãããšå¬ããã§ãïŒ
https://twitter.com/alexweberk
ä»åã® Colab ã¯ãã¡ãã§ãïŒ