ã¯ããã«
GoogleãããªãªãŒã¹ãããGeminiã®æ§èœè©äŸ¡ã«äœ¿ãããŠãããã³ãããŒã¯ã®æŠèŠããŸãšããŠã¿ãŸãããäžèšã®ããšãæåŸ ããŠããŸãã
ã»çŸåšã®AIãã©ãããããšãã©ã®çšåºŠã§ããããç¥ãã
ã»çŸåšã®AIãã©ã®ãããªããšã«åŒ±ãããç¥ãã
ã»æ°ããå€§èŠæš¡AIã¢ãã«ãç»å Žãããšãã«ãåªå£ãæ¯èŒã§ããããã«ããã
ã»ã¢ãã«ã«ãã£ãŠåŸæäžåŸæãããã®ã§ãè€æ°ã®AIãçšéã«å¿ããŠäœ¿ãåããããããã«ããã
ããã¹ã
äžè¬
MMLU
MMLU (Massive Multitask Language Understanding) ã¯2020幎ã«Center for AI Safety(CAIS)ã®Dan Hendrycksãã«ãã£ãŠææ¡ãããèšèªã¢ãã«ãè©äŸ¡ããããã®ãã³ãããŒã¯ã§ããåçæ°åŠãç±³åœå²ãã³ã³ãã¥ãŒã¿ ãµã€ãšã³ã¹ãæ³åŸãªã©ã®57ã®ç§ç®ããã multiple-choice tasks(å€è¢éžææ³ã®ã¿ã¹ã¯)ã§åé¡ãè§£ããŸãã
Measuring Massive Multitask Language Understanding
Abstract:
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability. We find that while most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average. However, on every one of the 57 tasks, the best models still need substantial improvements before they can reach expert-level accuracy. Models also have lopsided performance and frequently do not know when they are wrong. Worse, they still have near-random accuracy on some socially important subjects such as morality and law. By comprehensively evaluating the breadth and depth of a model's academic and professional understanding, our test can be used to analyze models across many tasks and to identify important shortcomings.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
ããã¹ãã¢ãã«ã®ãã«ãã¿ã¹ã¯ç²ŸåºŠã枬å®ããããã®æ°ãããã¹ããææ¡ããŸãããã®ãã¹ãã§ã¯ãåçæ°åŠãç±³åœå²ãã³ã³ãã¥ãŒã¿ ãµã€ãšã³ã¹ãæ³åŸãªã©ãå«ã 57 ã®èª²é¡ãåãäžããããŸãããã®ãã¹ãã§é«ã粟床ãéæããã«ã¯ãã¢ãã«ã¯äžçã«é¢ããåºç¯ãªç¥èãšåé¡è§£æ±ºèœåãåããŠããå¿ èŠããããŸããææ°ã®ã¢ãã«ã®ç²ŸåºŠã¯ã©ã³ãã ã«è¿ã粟床ã§ãããæå€§ã® GPT-3 ã¢ãã«ã¯ã©ã³ãã ã«æ¯ã¹ãŠå¹³åã§ã»ãŒ 20 ããŒã»ã³ãåäžããŠããããšãããããŸããããã ãã57 ã®ã¿ã¹ã¯ã®ããããã«ãããŠãæè¯ã®ã¢ãã«ããšãã¹ããŒã ã¬ãã«ã®ç²ŸåºŠã«éããã«ã¯ãäŸç¶ãšããŠå€§å¹ ãªæ¹åãå¿ èŠã§ããã¢ãã«ã®ããã©ãŒãã³ã¹ã«ãåããããããã€ééã£ãŠããã®ãããããªãããšããããããŸããããã«æªãããšã«ãéåŸ³ãæ³åŸãªã©ã®ç€ŸäŒçã«éèŠãªäž»é¡ã«é¢ããŠã¯ãäŸç¶ãšããŠã»ãŒã©ã³ãã ãªæ£ç¢ºæ§ãæã£ãŠããŸããã¢ãã«ã®åŠè¡çããã³å°éççè§£ã®åºããšæ·±ããå æ¬çã«è©äŸ¡ããããšã§ãç§ãã¡ã®ãã¹ãã䜿çšããŠãå€ãã®ã¿ã¹ã¯ã«ããã£ãŠã¢ãã«ãåæããéèŠãªæ¬ ç¹ãç¹å®ã§ããŸãã
æšè«
Big-Bench Hard(BBH)
BIG-bench(Beyond the Imitation Game benchmark)ã¯2022幎ã«Aarohi Srivastavaãã«ãã£ãŠææ¡ãããèšèªã¢ãã«ãè©äŸ¡ããããã®ãã³ãããŒã¯ã§ãã204ã®ã¿ã¹ã¯ã§æ§æãããŠããã132 æ©é¢ã® 450 人ã®èè ãã«ãã£ãŠäœããŠããŸããèšèªåŠãå¹Œå æã®çºéãæ°åŠãåžžèçæšè«ãçç©åŠãç©çåŠã瀟äŒçåèŠããœãããŠã§ã¢éçºãªã©ããåé¡ãäœãããŠããŸãã
BIG-Bench Hard (BBH) ã¯ã BIG-Benchã®äžã®23ã®å°é£ãªã¿ã¹ã¯ã§ãããããã¯ã以åã®èšèªã¢ãã«ã®è©äŸ¡ãå¹³åçãªäººéã®è©äŸ¡è ãäžåãææãäžããªãã£ãã¿ã¹ã¯ã§ã倿®µéã®æšè«ãèŠæ±ãããŸãã
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.
Abstract:
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG-bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit "breakthrough" behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
èšèªã¢ãã«ã¯ãèŠæš¡ã®å¢å ã«äŒŽãéçãªæ¹åãšæ°ãã質çãªæ©èœã®äž¡æ¹ãå®èšŒããŸããå€é©ãããããå¯èœæ§ã®ãã圱é¿ã«ããããããããããã®æ°ããæ©èœã¯ãŸã ååã«ç¹åŸŽä»ããããŠããŸãããå°æ¥ã®ç ç©¶ã«æ å ±ãæäŸããç Žå£çãªæ°ããã¢ãã«æ©èœã«åãã瀟äŒçæªåœ±é¿ã軜æžããã«ã¯ãçŸåšããã³è¿ãå°æ¥ã®èšèªã¢ãã«ã®æ©èœãšéçãçè§£ããããšãéèŠã§ãããã®èª²é¡ã«å¯ŸåŠããããã«ãBeyond the Imitation Game ãã³ãããŒã¯ (BIG ãã³ã) ãå°å ¥ããŸããBIG-bench ã¯çŸåš 204 ã®ã¿ã¹ã¯ã§æ§æãããŠããã132 æ©é¢ã® 450 人ã®èè ãå¯çš¿ããŠããŸããã¿ã¹ã¯ã®ãããã¯ã¯å€å²ã«ããããèšèªåŠãå¹Œå æã®çºéãæ°åŠãåžžèçæšè«ãçç©åŠãç©çåŠã瀟äŒçåèŠããœãããŠã§ã¢éçºãªã©ããåé¡ãæããŸããBIG-bench ã¯ãçŸåšã®èšèªã¢ãã«ã®èœåãè¶ ããŠãããšèããããã¿ã¹ã¯ã«çŠç¹ãåœãŠãŠããŸããOpenAI ã® GPT ã¢ãã«ãGoogle å éšã®ãã³ã¹ ãã©ã³ã¹ãã©ãŒã㌠ã¢ãŒããã¯ãã£ãããã³ã¹ã€ãã ã¹ã¿ã€ã«ã®ã¹ããŒã¹ ãã©ã³ã¹ãã©ãŒããŒã®åäœããæ°çŸäžããæ°ååã®ãã©ã¡ãŒã¿ã«ãããã¢ãã« ãµã€ãºã«ããã£ãŠ BIG ãã³ãã§è©äŸ¡ããŸããããã«ã匷åãªããŒã¹ã©ã€ã³ãæäŸããããã«ã人éã®å°éè©äŸ¡è ã®ããŒã ããã¹ãŠã®ã¿ã¹ã¯ãå®è¡ããŸããã調æ»çµæã«ã¯æ¬¡ã®ãã®ãå«ãŸããŸããã¢ãã«ã®ããã©ãŒãã³ã¹ãšãã£ãªãã¬ãŒã·ã§ã³ã¯ã©ã¡ããã¹ã±ãŒã«ã«å¿ããŠåäžããŸããã絶察çãªèгç¹ã§ (è©äŸ¡è ã®ããã©ãŒãã³ã¹ãšæ¯èŒãããš) å£ã£ãŠããŸããããã©ãŒãã³ã¹ã¯ã¢ãã« ã¯ã©ã¹éã§é©ãã»ã©äŒŒãŠããŸãããã¹ããŒã¹æ§ã«ããå©ç¹ããããŸããåŸã ã«ãã€äºæž¬éãã«æ¹åããã¿ã¹ã¯ã«ã¯ãéåžžãå€§èŠæš¡ãªç¥èãæèšã³ã³ããŒãã³ããå«ãŸããŸãããéèŠãªã¹ã±ãŒã«ã§ãç»æçãªãåäœã瀺ãã¿ã¹ã¯ã«ã¯ãå€ãã®å Žåãè€æ°ã®ã¹ããããã³ã³ããŒãã³ãããŸãã¯èåŒ±ãªææšãå«ãŸããŸãã瀟äŒçåèŠã¯éåžžãç¶æ³ãææ§ãªç°å¢ã§ã¯èŠæš¡ã倧ãããªãã«ã€ããŠå¢å ããŸãããããã¯ããã³ãããæç€ºããããšã§æ¹åã§ããŸãã
DROP
DROP(Discrete Reasoning Over the content of Paragraphs)ã¯2019幎ã«Duaãã«ãã£ãŠææ¡ãããæšè«ã®èªè§£åãè©äŸ¡ãããã³ãããŒã¯ã§ããå顿°ã¯55,000åã§æ®µèœã®å 容ãããå æ¬çã«çè§£ããŠæšè«ããèœåãè©äŸ¡ãããŸãã è«æå·çæç¹ã®SOTA(state-of-the-art)ã®ææ³ã§38.4%(F1ã¹ã³ã¢)ã®ç²ŸåºŠããéæã§ããªãã£ãããã§ã人éã®å°éå®¶ã®96%(F1ã¹ã³ã¢)ãããæšè«èœåããšãŠãäœããèšèªã¢ãã«ãæšè«ã«åŒ±ãã£ãããšã瀺ããŠããŸãã(Gemini Ultraã®DROPã®F1ã¹ã³ã¢ã¯82.4)
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.
Abstract:
Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 55k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs, as they remove the paraphrase-and-entity-typing shortcuts available in prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literatures on this dataset and show that the best systems only achieve 38.4% F1 on our generalized accuracy metric, while expert human performance is 96%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 51% F1.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
èªè§£åã¯æè¿æ¥éã«é²æ©ããŠãããã·ã¹ãã ã¯ã¿ã¹ã¯ã«æãäžè¬çãªããŒã¿ã»ããã人éãšç §åããããã«ãªããŸãããããããå€ãã®ç ç©¶ã«ãã£ãŠãããã®ã·ã¹ãã ã®è匱æ§ãæµ®ã圫ãã«ãªããããã¹ãããšããŸã ããããããããšã瀺ãããŠããŸããæ°ããèªè§£ãã³ãããŒã¯ã§ãã DROP ãå°å ¥ããŸããããã¯ã段èœã®å 容ã«å¯Ÿãã颿£æšè«ãå¿ èŠãšããŸãããã®ã¯ã©ãŠããœãŒã¹ã§æµå¯Ÿè ãäœæãã 55,000 åã®ãã³ãããŒã¯ã§ã¯ãã·ã¹ãã ã¯è³ªåå ã®åç § (ããããè€æ°ã®å ¥åäœçœ®) ã解決ãããããã«å¯ŸããŠåå¥ã®æäœ (å ç®ãã«ãŠã³ããäžŠã¹æ¿ããªã©) ãå®è¡ããå¿ èŠããããŸãããããã®æäœã§ã¯ã以åã®ããŒã¿ã»ããã§å©çšã§ããèšãæãããšã³ãã£ãã£ã®å ¥åã®ã·ã§ãŒãã«ãããåé€ããããããæ®µèœã®å 容ãããå æ¬çã«çè§£ããå¿ èŠããããŸãããã®ããŒã¿ã»ããã«å¯ŸããŠèªè§£ãšæå³è§£æã®äž¡æ¹ã®æç®ããåŸãæå ç«¯ã®ææ³ãé©çšããæè¯ã®ã·ã¹ãã ã¯äžè¬åãããç²ŸåºŠææšã§ 38.4% ã® F1 ããéæã§ããªãã®ã«å¯Ÿããå°éå®¶ã®äººéã®ããã©ãŒãã³ã¹ã¯ 96% ã§ããããšã瀺ããŸãããããã«ã51% ã® F1 ãéæããããã«ãèªè§£æ¹æ³ãšåçŽãªæ°çæšè«ãçµã¿åãããæ°ããã¢ãã«ãæç€ºããŸãã
HellaSwag
HellaSwagã¯ãOpenAIã®ç ç©¶å¡ã§ããRowan Zellersãã«ãã£ãŠ2019å¹Žã«ææ¡ãããèªç¶èšèªã®åžžèçãªæšè«ãåããã³ãããŒã¯ã§ãã人éã«ã¯ããããªè³ªå(粟床ã95%以äž)ã§ãã£ãŠããæå 端ã¢ãã«ã§ã¯åçãå°é£ãªããŒã¿ã»ãã(粟床ã48ïŒ æªæº)ãšãªã£ãŠããŸãã
Hellaswag: Can a machine really finish your sentence?
Abstract:
Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the keys." With the introduction of BERT, near human-level performance was reached. Does this mean that machines can perform human level commonsense inference?
In this paper, we show that commonsense inference still proves difficult for even state-of-the-art models, by presenting HellaSwag, a new challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). We achieve this via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly robust. The key insight is to scale up the length and complexity of the dataset examples towards a critical 'Goldilocks' zone wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models.
Our construction of HellaSwag, and its resulting difficulty, sheds light on the inner workings of deep pretrained models. More broadly, it suggests a new path forward for NLP research, in which benchmarks co-evolve with the evolving state-of-the-art in an adversarial way, so as to present ever-harder challenges.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
Zellersãã«ããæè¿ã®ç ç©¶ã(2018) ã¯ãåžžèçãªèªç¶èšèªæšè«ã®æ°ããã¿ã¹ã¯ãå°å ¥ããŸãããã女æ§ããã¢ãã®åã«åº§ã£ãŠããããªã©ã®ã€ãã³ãã®èª¬æãäžãããããšããã·ã³ã¯æãå¯èœæ§ã®é«ãåŸç¶ãéžæããªããã°ãªããŸãã:ã圌女ã¯éµç€ã«æã眮ãããBERT ã®å°å ¥ã«ãããã»ãŒäººéã¬ãã«ã®ããã©ãŒãã³ã¹ã«å°éããŸãããããã¯ãæ©æ¢°ã人éã¬ãã«ã®åžžèçãªæšè«ãå®è¡ã§ããããšãæå³ããŸãã?
ãã®è«æã§ã¯ãæ°ãã課é¡ããŒã¿ã»ããã§ãã HellaSwag ãæç€ºããããšã«ãããåžžèçãªæšè«ã¯æå 端ã®ã¢ãã«ã§ãäŸç¶ãšããŠé£ããããšã倿ããŠããããšã瀺ããŸãããã®è³ªåã¯äººéã«ãšã£ãŠã¯äºçްãªãã®ã§ãã (粟床ã 95% 以äž)ãæå 端ã®ã¢ãã«ã¯å°é£ã䌎ããŸã (粟床ã 48% æªæº)ãããã¯ãäžé£ã®èå¥åãæ©æ¢°ã«ãã£ãŠçæãããæµå¯Ÿçãªééã£ãåçã®ã»ãããç¹°ãè¿ãéžæããããŒã¿åéãã©ãã€ã ã§ããæµå¯Ÿçãã£ã«ã¿ãªã³ã° (AF) ã«ãã£ãŠå®çŸãããŸããAFã¯é©ãã»ã©å ç¢ã§ããããšãããããŸããéèŠãªæŽå¯ã¯ãçæãããããã¹ãã人éã«ãšã£ãŠã°ãã°ãããã«ãããããããæå 端ã®ã¢ãã«ã«ãã£ãŠèª€åé¡ãããããšãå€ãã¯ãªãã£ã«ã«ãªããŽã«ãã£ããã¯ã¹ããŸãŒã³ã«åããŠãããŒã¿ã»ããã®äŸã®é·ããšè€éããã¹ã±ãŒã«ã¢ããããããšã§ãã
ç§ãã¡ã® HellaSwag ã®æ§ç¯ãšãã®çµæãšããŠçããå°é£ãã¯ãæ·±ãäºåãã¬ãŒãã³ã°ãããã¢ãã«ã®å éšåäœã«å ãåœãŠãŸããããåºç¯ã«ã¯ããã³ãããŒã¯ãæµå¯Ÿçãªæ¹æ³ã§é²åããæå 端æè¡ãšå ±é²åãããããŸã§ä»¥äžã«å°é£ãªèª²é¡ãæç€ºãããNLP ç ç©¶ã®æ°ããªåé²ã®éã瀺åããŠããŸãã
Math
GSM8K
GSM8Kã¯ãOpenAIã®ãªãµãŒããµã€ãšã³ãã£ã¹ãã§ããKarl Cobbeãã«ãã£ãŠã2021å¹Žã«ææ¡ãããè€æ°ã¹ãããã®æ°åŠçæšè«ã®ãã³ãããŒã¯ã§ãã8.5Kã®é«å質ã§èšèªçã«å€æ§ãªå°åŠæ ¡ã®æ°åŠã®æç« åé¡ã®ããŒã¿ã»ããã䜿ã£ãŠããŸããèè ãã«ãããšçŸåšã®ã¢ãã«ã¯ãã®è€æ°ã¹ãããã®æ°åŠçæšè«ã匱ã¿ã§ãããšã®ããšã§ãã
Training verifiers to solve math word problems
Abstract:
State-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning. To diagnose the failures of current models and support research, we introduce GSM8K, a dataset of 8.5K high quality linguistically diverse grade school math word problems. We find that even the largest transformer models fail to achieve high test performance, despite the conceptual simplicity of this problem distribution. To increase performance, we propose training verifiers to judge the correctness of model completions. At test time, we generate many candidate solutions and select the one ranked highest by the verifier. We demonstrate that verification significantly improves performance on GSM8K, and we provide strong empirical evidence that verification scales more effectively with increased data than a finetuning baseline.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
æå 端ã®èšèªã¢ãã«ã¯ãå€ãã®ã¿ã¹ã¯ã§äººéã®ããã©ãŒãã³ã¹ã«å¹æµããããšãã§ããŸãããè€æ°ã¹ãããã®æ°åŠçæšè«ã確å®ã«å®è¡ããã®ã¯ãŸã å°é£ã§ããçŸåšã®ã¢ãã«ã®é害ã蚺æããç ç©¶ããµããŒãããããã«ã8.5K ã®é«å質ã§èšèªçã«å€æ§ãªå°åŠæ ¡ã®æ°åŠã®æç« åé¡ã®ããŒã¿ã»ããã§ãã GSM8K ãå°å ¥ããŸãããã®åé¡ååžã®æŠå¿µçãªåçŽãã«ãããããããæå€§ã®å€å§åšã¢ãã«ã§ãé«ããã¹ãæ§èœãéæã§ããªãããšãããããŸãããããã©ãŒãã³ã¹ãåäžãããããã«ãã¢ãã«ã®å®æåºŠã®æ£ç¢ºãã倿ãããã¬ãŒãã³ã°æ€èšŒè ãææ¡ããŸãããã¹ãæã«ã¯ãå€ãã®åè£è§£ãçæãããæ€èšŒè ã«ãã£ãŠæãé«ãã©ã³ã¯ãä»ãããããã®ãéžæãããŸããæ€èšŒã«ãã£ãŠ GSM8K ã®ããã©ãŒãã³ã¹ãå€§å¹ ã«åäžããããšãå®èšŒãã埮調æŽããŒã¹ã©ã€ã³ãããããŒã¿ã®å¢å ã«å¿ããŠæ€èšŒããã广çã«æ¡åŒµããããšãã匷åãªçµéšç蚌æ ãæäŸããŸãã
MATH
MATHã¯Center for AI Safety(CAIS)ã®ãã£ã¬ã¯ã¿ãŒã§ããDan Hendrycksãã«ãã£ãŠã2021å¹Žã«ææ¡ããããã³ãããŒã¯ã§ãã12,500ã®ãã£ã¬ã³ãžã³ã°ãªæ°åŠçåé¡ãããªãããŒã¿ã»ããã䜿ã£ãŠããŠãååé¡ã«ã¯å®å
šãªã¹ããããã€ã¹ãããã®è§£æ±ºçããããŸãã
èè
ãã«ãããšã巚倧ãªTransformerã¢ãã«ã§ãã£ãŠãç²ŸåºŠãæ¯èŒçäœããŸãŸã§ãããã¢ãã«ã®ãã©ã¡ãŒã¿ãŒæ°ãå¢ããã ãã§ã¯ã匷åçãªæ°åŠçæšè«ãéæããã®ã¯éçŸå®çã§ãããšã®ããšã§ãã
Measuring mathematical problem solving with the MATH dataset.
Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations. To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics. Even though we are able to increase accuracy on MATH, our results show that accuracy remains relatively low, even with enormous Transformer models. Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue. While scaling Transformers is automatically solving most other text-based tasks, scaling is not currently solving MATH. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community.
ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
å€ãã®ç¥çäœæ¥ã«ã¯æ°åŠçãªåé¡è§£æ±ºãå¿ èŠã§ããããã®ã¹ãã«ã¯äŸç¶ãšããŠã³ã³ãã¥ãŒã¿ãŒã®èœåãè¶ ããŠããŸããæ©æ¢°åŠç¿ã¢ãã«ã§ãã®èœåãæž¬å®ããããã«ã12,500 ã®ææŠçãªæ°åŠã®åé¡ã®æ°ããããŒã¿ã»ããã§ãã MATH ãå°å ¥ããŸããMATH ã®ååé¡ã«ã¯å®å šãªã¹ããããã€ã¹ãããã®è§£æ±ºçããããã¢ãã«ã«çãã®å°åºãšèª¬æãçæããããæããããã«äœ¿çšã§ããŸããå°æ¥ã®ç ç©¶ãä¿é²ããæ°åŠã®ç²ŸåºŠãåäžãããããã«ãã¢ãã«ã«æ°åŠã®åºç€ãæããã®ã«åœ¹ç«ã€å€§èŠæš¡ãªè£å©äºåãã¬ãŒãã³ã° ããŒã¿ã»ãããæäŸããŠããŸããMATH ã®ç²ŸåºŠãåäžãããããšã¯ã§ããŸãããã巚倧㪠Transformer ã¢ãã«ã§ãã£ãŠãç²ŸåºŠãæ¯èŒçäœããŸãŸã§ããããšãçµæããããããŸããããã«ãã¹ã±ãŒãªã³ã°ã®åŸåãç¶ãå Žåãåã«äºç®ãšã¢ãã«ã®ãã©ã¡ãŒã¿ãŒæ°ãå¢ããã ãã§ã¯ã匷åãªæ°åŠçæšè«ãéæããã®ã¯éçŸå®çã§ããããšãããããŸãããTransformers ã®ã¹ã±ãŒãªã³ã°ã¯ä»ã®ã»ãšãã©ã®ããã¹ãããŒã¹ã®ã¿ã¹ã¯ãèªåçã«è§£æ±ºããŸãããã¹ã±ãŒãªã³ã°ã¯çŸåš MATH ã解決ããŸãããæ°åŠçåé¡è§£æ±ºãããã«æšé²ããã«ã¯ãããåºç¯ãªç ç©¶ã³ãã¥ããã£ããã®æ°ããã¢ã«ãŽãªãºã ã®é²æ©ãå¿ èŠã«ãªãã§ãããã
Code
HumanEval
HumanEvalã¯OpenAIã®ãªãµãŒããµã€ãšã³ãã£ã¹ãã§ããMark Chenãã«ãã£ãŠã2021å¹Žã«ææ¡ãããæååããããã°ã©ã ãçæããæ©èœã®æ£ãããæž¬å®ãããã³ãããŒã¯ã§ãã
èè
ãã®ã¢ãã«Codex(GitHubã³ãŒãã§ãã¡ã€ã³ãã¥ãŒãã³ã°ããGPTèšèªã¢ãã«)ã¯åé¡ã®28.8%ã解決ããŠãGPT-3ã¯0%ãGPT-Jã¯11.4%ã®åé¡ã解決ããããã§ãããŸãã¢ãã«ããã®ãµã³ããªã³ã°ãç¹°ãè¿ãããšã§èè
ãã®ã¢ãã«ã¯70.2%ã®åé¡ã解決ããããã§ãã
Evaluating large language models trained on code
Abstract:
We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot. On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems, while GPT-3 solves 0% and GPT-J solves 11.4%. Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts. Using this method, we solve 70.2% of our problems with 100 samples per problem. Careful investigation of our model reveals its limitations, including difficulty with docstrings describing long chains of operations and with binding operations to variables. Finally, we discuss the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
GitHub ããå ¬éãããŠããã³ãŒãã«åºã¥ããŠåŸ®èª¿æŽããã GPT èšèªã¢ãã«ã§ãã Codex ã玹ä»ãããã® Python ã³ãŒãäœææ©èœãç ç©¶ããŸããCodex ã®ç¬èªã®è£œåããŒãžã§ã³ã GitHub Copilot ã匷åããŸããããã¥ã¡ã³ãæååããããã°ã©ã ãåæããæ©èœã®æ£ãããæž¬å®ããããã«ãªãªãŒã¹ãããæ°ããè©äŸ¡ã»ããã§ãã HumanEval ã§ã¯ãç§ãã¡ã®ã¢ãã«ã¯åé¡ã® 28.8% ã解決ããŸããããGPT-3 㯠0%ãGPT-J 㯠11.4% ã解決ããŸãããããã«ãã¢ãã«ããã®ãµã³ããªã³ã°ãç¹°ãè¿ãããšããå°é£ãªããã³ããã«å¯ŸããŠæå¹ãªè§£æ±ºçãçã¿åºãããã®é©ãã»ã©å¹æçãªæŠç¥ã§ããããšãããããŸããããã®æ¹æ³ã䜿çšãããšãåé¡ããšã« 100 åã®ãµã³ãã«ã䜿çšããŠåé¡ã® 70.2% ã解決ã§ããŸããç§ãã¡ã®ã¢ãã«ãæ³šææ·±ã調æ»ãããšãé·ãæäœãã§ãŒã³ãèšè¿°ããããã¥ã¡ã³ãæååã倿°ãžã®æäœã®ãã€ã³ãã®é£ãããªã©ããã®éçãæããã«ãªããŸããæåŸã«ãå®å šæ§ãã»ãã¥ãªãã£ãçµæžæ§ãã«ããŒããã匷åãªã³ãŒãçæãã¯ãããžã®å°å ¥ã«ããæœåšçãªåºç¯ãªåœ±é¿ã«ã€ããŠèª¬æããŸãã
MULTIMODAL
Image Understanding
MMMU
MMMUã¯ãªãã€ãªå·ç«å€§åŠã®Xiang Yueãã«ãã£ãŠã2023幎ã«èæ¡ããã倧åŠã¬ãã«ã®ç¥èãšæšè«ã«é¢ãããã³ãããŒã¯ã§ãã倧åŠã®è©Šéšãã¯ã€ãºãæç§æžããæ³šææ·±ãåéããã 11.5Kã®ãã«ãã¢ãŒãã«ãªè³ªåã§æ§æãããŠããŸããæ±çšäººå·¥ç¥èœïŒAGIïŒã®ã¬ãã«3ãšããŠå®çŸ©ãããããšãã¹ããŒãAGIãã®é²æ©ãè©äŸ¡ãããã³ãããŒã¯ãšããŠæçšæ§ãããããã§ãã
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Abstract:
We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly heterogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. Our evaluation of 14 open-source LMMs and the proprietary GPT-4V(ision) highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V only achieves a 56% accuracy, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
MMMUã玹ä»ããŸããMMMUã¯ã倧åŠã¬ãã«ã®äž»é¡ç¥èãšæå³çãªæšè«ãå¿ èŠãšããå€§èŠæš¡ãªè€æ°åéã®ã¿ã¹ã¯ã§ãã«ãã¢ãŒãã«ã¢ãã«ãè©äŸ¡ããããã«èšèšãããæ°ãããã³ãããŒã¯ã§ããMMMUã«ã¯ãèžè¡ãšãã¶ã€ã³ãããžãã¹ãç§åŠãå¥åº·ãšå»åŠã人æç§åŠãšç€ŸäŒç§åŠãæè¡ãšå·¥åŠã®6ã€ã®äž»èŠåéãã«ããŒããã倧åŠã®è©Šéšãã¯ã€ãºãæç§æžããæ³šææ·±ãåéããã11.5Kã®ãã«ãã¢ãŒãã«ãªè³ªåãå«ãŸããŠããŸãããããã®è³ªåã¯ã30ã®äž»é¡ãš183ã®ãµããã£ãŒã«ãã«åã³ããã£ãŒããå³ãå°å³ãè¡šãæ¥œèãååŠæ§é ãªã©ã30çš®é¡ã®éåžžã«ç°è³ªãªç»åã§æ§æãããŠããŸããæ¢åã®ãã³ãããŒã¯ãšã¯ç°ãªããMMMUã¯ãå°éå®¶ãçŽé¢ããã¿ã¹ã¯ãšåæ§ã®ã¿ã¹ã¯ãå®è¡ããããã®ããã¡ã€ã³åºæã®ç¥èã«ããé«åºŠãªèªèãšæšè«ãææŠçãªã¢ãã«ã«çŠç¹ãåœãŠãŠããŸãã14ã®ãªãŒãã³ãœãŒã¹LMMãšç¬èªã®GPT-4V(ision)ã®è©äŸ¡ã§ã¯ãMMMU ã«ãã£ãŠããããããé倧ãªèª²é¡ãæµ®ã圫ãã«ãªããŸãããå é²ç㪠GPT-4Vã§ãã56%ã®ç²ŸåºŠããéæã§ãããæ¹åã®äœå°ã倧ããããšãããããŸããç§ãã¡ã¯ãMMMUãã³ãã¥ããã£ãåºæ¿ããŠããšãã¹ããŒãã®æ±çšäººå·¥ç¥èœã«åããæ¬¡äžä»£ã®ãã«ãã¢ãŒãã«åºç€ã¢ãã«ãæ§ç¯ãããšä¿¡ããŠããŸãã
VQAv2
VQAv2ã¯ããŒãžãã¢å·¥ç§å€§åŠã®Yash Goyalãã«ãã£ãŠã2017幎ã«èæ¡ãããç»åçè§£ã«é¢ãããã³ãããŒã¯ã§ãã1ã€ã®è³ªåã«å¯ŸããŠïŒã€ã®ç»åããããããªè³ªåã®ããŒã¿ã»ãããšãªã£ãŠãããããžã¥ã¢ã«è³ªåå¿ç (VQA) ã¿ã¹ã¯ã®VãéèŠãããã³ãããŒã¯ãšãªã£ãŠããŸãã
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Abstract:
Problems at the intersection of vision and language are of significant importance both as challenging research questions and for the rich set of applications they enable. However, inherent structure in our world and bias in our language tend to be a simpler signal for learning than visual modalities, resulting in models that ignore visual information, leading to an inflated sense of their capability.
We propose to counter these language priors for the task of Visual Question Answering (VQA) and make vision (the V in VQA) matter! Specifically, we balance the popular VQA dataset by collecting complementary images such that every question in our balanced dataset is associated with not just a single image, but rather a pair of similar images that result in two different answers to the question. Our dataset is by construction more balanced than the original VQA dataset and has approximately twice the number of image-question pairs. Our complete balanced dataset is publicly available at this http URL as part of the 2nd iteration of the Visual Question Answering Dataset and Challenge (VQA v2.0).
We further benchmark a number of state-of-art VQA models on our balanced dataset. All models perform significantly worse on our balanced dataset, suggesting that these models have indeed learned to exploit language priors. This finding provides the first concrete empirical evidence for what seems to be a qualitative sense among practitioners.
Finally, our data collection protocol for identifying complementary images enables us to develop a novel interpretable model, which in addition to providing an answer to the given (image, question) pair, also provides a counter-example based explanation. Specifically, it identifies an image that is similar to the original image, but it believes has a different answer to the same question. This can help in building trust for machines among their users.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
èŠèŠãšèšèªã亀差ããåé¡ã¯ãç ç©¶äžã®å°é£ãªèª²é¡ãšããŠãããããå¯èœã«ããè±å¯ãªå¿çšãšããŠãéåžžã«éèŠã§ããããããç§ãã¡ã®äžçã«åºæã®æ§é ãèšèªã®åãã¯ãèŠèŠçãªã¢ããªãã£ãããåŠç¿ã®ããã®åçŽãªã·ã°ãã«ãšãªãåŸåãããããã®çµæãèŠèŠçãªæ å ±ãç¡èŠããã¢ãã«ãçæãããã¢ãã«ã®èœåãèªåŒµãããæèŠã«ã€ãªãããŸãã
ç§ãã¡ã¯ãããžã¥ã¢ã«è³ªåå¿ç (VQA) ã®ã¿ã¹ã¯ã«é¢ãããããã®èšèªã®äºåæ¡ä»¶ã«å¯Ÿæããããžã§ã³ (VQA ã® V) ãéèŠãªãã®ã«ããããšãææ¡ããŸããå ·äœçã«ã¯ããã©ã³ã¹ã®ãšããããŒã¿ã»ããå ã®ãã¹ãŠã®è³ªåã 1 ã€ã®ç»åã ãã§ã¯ãªãã質åã«å¯Ÿãã 2 ã€ã®ç°ãªãåçãããããé¡äŒŒããç»åã®ãã¢ã«é¢é£ä»ããããããã«ãçžè£çãªç»åãåéããããšã§ã人æ°ã®ãã VQA ããŒã¿ã»ããã®ãã©ã³ã¹ããšããŸããç§ãã¡ã®ããŒã¿ã»ããã¯ãå ã® VQA ããŒã¿ã»ããããããã©ã³ã¹ã®åããæ§é ã«ãªã£ãŠãããç»åãšè³ªåã®ãã¢ã®æ°ãçŽ 2 åã«ãªã£ãŠããŸããç§ãã¡ã®å®å šãªãã©ã³ã¹ã®åããããŒã¿ã»ããã¯ãVisual Question Answering Dataset and Challenge (VQA v2.0) ã® 2 åç®ã®å埩ã®äžéšãšããŠã ãã® http URLã§å ¬éãããŠããŸãã
ããã«ããã©ã³ã¹ã®åããããŒã¿ã»ããã§å€æ°ã®æå 端㮠VQA ã¢ãã«ããã³ãããŒã¯ããŸãããã©ã³ã¹ã®ãšããããŒã¿ã»ããã§ã¯ãã¹ãŠã®ã¢ãã«ã®ããã©ãŒãã³ã¹ãå€§å¹ ã«äœäžããŠããããããã®ã¢ãã«ãå®éã«èšèªäºåååžãå©çšããããšãåŠç¿ããŠããããšã瀺åããŠããŸãããã®çºèŠã¯ãå®è·µè ã®éã§å®æ§çæèŠãšæããããã®ã«å¯ŸããåããŠã®å ·äœçãªçµéšç蚌æ ãæäŸãããã®ã§ããã
æåŸã«ãçžè£çãªç»åãèå¥ããããã®ããŒã¿åéãããã³ã«ã«ãããäžãããã (ç»åã質å) ãã¢ã«å¯ŸããçããæäŸããã ãã§ãªããåäŸã«åºã¥ãã説æãæäŸãããæ°ããè§£éå¯èœãªã¢ãã«ãéçºããããšãã§ããŸããå ·äœçã«ã¯ãå ã®ç»åã«äŒŒãŠããããåã質åã«å¯ŸããŠç°ãªãçãããããšèããããç»åãèå¥ããŸããããã¯ããŠãŒã¶ãŒéã§ãã·ã³ã«å¯Ÿããä¿¡é Œãæ§ç¯ããã®ã«åœ¹ç«ã¡ãŸãã
TextVQA
TextVQAã¯Facebook AI Researchã®Amanpreet Singhãã«ãã£ãŠã2019幎ã«èæ¡ãããç»åå ã«æžãããããã¹ããèªã¿åãæšè«ããèœåãæž¬ããã³ãããŒã¯ã§ããçŸåšã®VQAã¢ãã«ã§ã¯ãç»åå ã®ããã¹ããèªã¿åã£ãŠæšè«ããŠçããããšãé£ããããã§ãVQAv2ãè£å®ãããã³ãããŒã¯ãšãªãããã§ãã
Towards VQA models that can read
Abstract:
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
ç ç©¶ã«ãããšãèŠèŠé害ã®ãããŠãŒã¶ãŒãåšå²ã®ç»åã«é¢ããŠè¡ã質åã®äž»ãªçš®é¡ã¯ãç»åå ã®ããã¹ããèªãããšã«é¢ãããã®ã§ããããšãããã£ãŠããŸããããããä»ã® VQA ã¢ãã«ã¯èªã¿åããŸãããç§ãã¡ã®è«æã¯ããã®åé¡ã«å¯ŸåŠããããã®ç¬¬äžæ©ãèžã¿åºããŸãããŸãããã®éèŠãªåé¡ã®é²æãä¿é²ããããã«ãæ°ãããTextVQAãããŒã¿ã»ãããå°å ¥ããŸããæ¢åã®ããŒã¿ã»ããã«ã¯ãããã¹ãã«é¢ãã質åã®å²åãå°ãªãã (VQA ããŒã¿ã»ãããªã©)ãå°ããããŸã (VizWiz ããŒã¿ã»ãããªã©)ãTextVQA ã«ã¯ã28,408 æã®ç»åã«é¢ãã 45,336 åã®è³ªåãå«ãŸããŠãããåçããã«ã¯ããã¹ãã«ã€ããŠã®æšè«ãå¿ èŠã§ãã2 çªç®ã«ãç»åå ã®ããã¹ããèªã¿åããç»åãšè³ªåã®ã³ã³ããã¹ãã§ããã«ã€ããŠæšè«ããããã¹ããšç»åã«åºã¥ãããŸãã¯èŠã€ãã£ãæååã§æ§æãããæšè«ã§ããå¯èœæ§ã®ããçããäºæž¬ãããæ°ããã¢ãã« ã¢ãŒããã¯ãã£ãå°å ¥ããŸããç»åã§ã¯ããããã£ãŠãç§ãã¡ã¯ãã®ã¢ãããŒãã LookãReadãReason & Answer (LoRRA) ãšåŒãã§ããŸããLoRRA ã TextVQA ããŒã¿ã»ããäžã®æ¢åã®æå 端㮠VQA ã¢ãã«ãããåªããããã©ãŒãã³ã¹ãçºæ®ããããšã瀺ããŸããTextVQA ã§ã¯äººéã®ããã©ãŒãã³ã¹ãšãã·ã³ã®ããã©ãŒãã³ã¹ã®å·®ã VQA 2.0 ãããå€§å¹ ã«å€§ããããšãããããTextVQA ã VQA 2.0 ãè£å®ããæ¹åã«æ²¿ã£ã鲿©ã®ãã³ãããŒã¯ã«é©ããŠããããšã瀺åããŠããŸãã
DocVQA
DocVQAã¯IIIT Hyderabadã®Minesh Mathewãã«ãã£ãŠã2021幎ã«èæ¡ãããææžç»åã®çè§£ã«é¢ãããã³ãããŒã¯ã§ããå³ããã€ã¢ã°ã©ã ãã€ã³ãã©ã°ã©ãã£ãã¯çãèšèŒãããææžç»åã®æ å ±ãèŠèŠçã«çè§£ããããšãèŠæ±ãããŸãã
DocVQA: A Dataset for VQA on Document Images
Abstract:
We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at this http URLã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
DocVQA ãšåŒã°ãããããã¥ã¡ã³ãç»åäžã® Visual Question Answering (VQA) çšã®æ°ããããŒã¿ã»ããã玹ä»ããŸãããã®ããŒã¿ã»ããã¯ã12,000 以äžã®ããã¥ã¡ã³ãç»åã«å®çŸ©ããã 50,000 ã®è³ªåã§æ§æãããŠããŸããVQA ããã³èªè§£åã«é¢ããåæ§ã®ããŒã¿ã»ãããšæ¯èŒããããŒã¿ã»ããã®è©³çްãªåæã瀺ãããŠããŸããæ¢åã® VQA ãšèªè§£ã¢ãã«ãæ¡çšããŠãããã€ãã®ããŒã¹ã©ã€ã³çµæãå ±åããŸããæ¢åã®ã¢ãã«ã¯ãç¹å®ã®çš®é¡ã®è³ªåã§ã¯ããªãåªããããã©ãŒãã³ã¹ãçºæ®ããŸããã人éã®ããã©ãŒãã³ã¹ (粟床 94.36%) ãšæ¯èŒãããšãããã©ãŒãã³ã¹ã«å€§ããªã®ã£ããããããŸããã¢ãã«ã¯ãææžã®æ§é ãçè§£ããããšãéèŠãªè³ªåã«é¢ããŠç¹ã«æ¹åããå¿ èŠããããŸããããŒã¿ã»ãããã³ãŒãããªãŒããŒããŒãã¯ããã® http URLããå ¥æã§ããŸãã
Infographic VQA
Infographic VQAã¯IIIT Hyderabadã®Minesh Mathewãã«ãã£ãŠã2022幎ã«èæ¡ãããã€ã³ãã©ã°ã©ãã£ãã¯ç»åã®çè§£ã«é¢ãããã³ãããŒã¯ã§ããã€ã³ãã©ã°ã©ãã£ãã¯ç»åã¯ãããã¹ããã°ã©ãã£ãã¯ãããžã¥ã¢ã«èŠçŽ ã®çµã¿åããã䜿çšããŠæ å ±ã广çã«äŒéããããã«èšèšãããããã¥ã¡ã³ãã§ãã
Abstract:
Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic understanding of infographic images by using Visual Question Answering this http URL this end, we present InfographicVQA, a new dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations. The collected questions require methods to jointly reason over the document layout, textual content, graphical elements, and data visualizations. We curate the dataset with emphasis on questions that require elementary reasoning and basic arithmetic skills. Finally, we evaluate two strong baselines based on state of the art multi-modal VQA models, and establish baseline performance for the new task. The dataset, code and leaderboard will be made available at this http URLã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
ã€ã³ãã©ã°ã©ãã£ãã¯ã¹ã¯ãããã¹ããã°ã©ãã£ãã¯ãããžã¥ã¢ã«èŠçŽ ã®çµã¿åããã䜿çšããŠæ å ±ã广çã«äŒéããããã«èšèšãããããã¥ã¡ã³ãã§ãããã®ç ç©¶ã§ã¯ããã® http URL ã®Visual Question Answering ã䜿çšããŠãã€ã³ãã©ã°ã©ãã£ãã¯ç»åã®èªåçè§£ãæ¢çŽ¢ããŸããæåŸã«ãèªç¶èšèªã®è³ªåãšåçã®æ³šéãšãšãã«ãã€ã³ãã©ã°ã©ãã£ãã¯ã®å€æ§ãªã³ã¬ã¯ã·ã§ã³ã§æ§æãããæ°ããããŒã¿ã»ããã§ãã InfographicVQA ã玹ä»ããŸããåéããã質åã«ã¯ãããã¥ã¡ã³ãã®ã¬ã€ã¢ãŠããããã¹ãã®å 容ãã°ã©ãã£ãã¯èŠçŽ ãããã³ããŒã¿ã®èŠèŠåãå ±åã§æšè«ããæ¹æ³ãå¿ èŠã§ããç§ãã¡ã¯ãåæ©çãªæšè«ãšåºæ¬çãªç®è¡ã¹ãã«ãå¿ èŠãšãã質åã«éç¹ã眮ããŠããŒã¿ã»ãããå³éžããŠããŸããæåŸã«ãæå 端ã®ãã«ãã¢ãŒãã« VQA ã¢ãã«ã«åºã¥ã㊠2 ã€ã®åŒ·åãªããŒã¹ã©ã€ã³ãè©äŸ¡ããæ°ããã¿ã¹ã¯ã®ããŒã¹ã©ã€ã³ ããã©ãŒãã³ã¹ã確ç«ããŸããããŒã¿ã»ãããã³ãŒãããªãŒããŒããŒãã¯ããã® http URLããå ¥æã§ããŸãã
MathVista
MathVistaã¯UCLA Computer Science Departmentã®Pan Luãã«ãã£ãŠã2023幎ã«èæ¡ãããèŠèŠçãªã³ã³ããã¹ãã§ã®æ°åŠçæšè«ãã³ãããŒã¯ã§ããMathVistaã®ã¿ã¹ã¯ãå®äºããã«ã¯ãã现ããæ·±ãèŠèŠççè§£ãšæ§æçæšè«ãå¿ èŠã§ãæå 端ã®åºç€ã¢ãã«ã¯ãã®ãã¹ãŠãå°é£ã§ãããšèè ãã¯è¿°ã¹ãŠããŸãã
Abstract:
Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at this https URL.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
å€§èŠæš¡èšèªã¢ãã« (LLM) ãšå€§èŠæš¡ãã«ãã¢ãŒãã« ã¢ãã« (LMM) ã¯ãå€ãã®ã¿ã¹ã¯ãé åã§åªããåé¡è§£æ±ºã¹ãã«ã瀺ããŸãããèŠèŠçãªã³ã³ããã¹ãã§ã®æ°åŠçæšè«ã«ãããèœåã¯äœç³»çã«ç ç©¶ãããŠããŸããããã®ã®ã£ãããåããããã«ãããŸããŸãªæ°åŠçã¿ã¹ã¯ãšèŠèŠçã¿ã¹ã¯ããã®èª²é¡ãçµã¿åãããããã«èšèšããããã³ãããŒã¯ã§ãã MathVista ã玹ä»ããŸããããã¯ãæ°åŠãå«ã 28 ã®æ¢åã®ãã«ãã¢ãŒãã« ããŒã¿ã»ãããšãæ°ããäœæããã 3 ã€ã®ããŒã¿ã»ãã (IQTestãFunctionQAãããã³ PaperQA) ããæŽŸçãã 6,141 ã®äŸã§æ§æãããŠããŸãããããã®ã¿ã¹ã¯ãå®äºããã«ã¯ããã现ããæ·±ãèŠèŠççè§£ãšæ§æçæšè«ãå¿ èŠã§ãããæå 端ã®åºç€ã¢ãã«ã¯ãã¹ãŠãããå°é£ã§ãããšæããŠããŸããMathVista ã䜿çšããŠã12 ã®èåãªåºç€ã¢ãã«ã®å æ¬çãã€å®éçãªè©äŸ¡ã宿œããŸãããæé«ã®ããã©ãŒãã³ã¹ãèªã GPT-4V ã¢ãã«ã¯ãå šäœã®ç²ŸåºŠ 49.9% ãéæãã2 çªç®ã«åªããããã©ãŒãã³ã¹ãèªã Bard ã 15.1% äžåã£ãŠããŸããç§ãã¡ã®è©³çްãªåæã«ãããGPT-4V ã®åªäœæ§ã¯äž»ã«èŠèŠèªèãšæ°åŠçæšè«ã®åŒ·åã«èµ·å ããããšãæããã«ãªããŸããããã ããGPT-4V ã¯è€éãªæ°å€ãçè§£ããå³å¯ãªæšè«ãå®è¡ããã®ã«èŠåŽããããšãå€ãããã人éã®ããã©ãŒãã³ã¹ã«ã¯ãŸã 10.4% åã°ãªãããã®å€§ããªã®ã£ããã¯ãæ°åŠçã«éäžçã§èŠèŠçã«è±å¯ãªçŸå®äžçã®ã¿ã¹ã¯ã«åãçµãããšãã§ããæ±çš AI ãšãŒãžã§ã³ãã®éçºã«ãããŠãMathVista ãæããéèŠãªåœ¹å²ã匷調ããŠããŸããããã«ãGPT-4V ã®æ°ããèªå·±æ€èšŒæ©èœãèªå·±äžè²«æ§ã®é©çšã察話åãã£ãããããæ©èœã調æ»ããå°æ¥ã®ç ç©¶ã«ãããææãªå¯èœæ§ã匷調ããŸãããããžã§ã¯ãã¯ããã® https URLã§å ¥æã§ããŸãã
Video Understanding
VATEX
VATEXã¯University of California, Santa Barbaraã®Xin Wangãã«ãã£ãŠã2019幎ã«èæ¡ãããå€èšèªãããªã«é¢ãããã³ãããŒã¯ã§ããVATEXã¯è±èªãšäžåœèªã®äž¡æ¹ã§41,250以äžã®ãããªãš825,000ã®ãã£ãã·ã§ã³ãå«ã¿ãå€èšèªå¯Ÿå¿ã§ãå€§èŠæš¡ã§ãèšèªçã«è€éã§ãããšèè ãã¯è¿°ã¹ãŠããŸãã
VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
Abstract:
We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSR-VTT dataset, VATEX is multilingual, larger, linguistically complex, and more diverse in terms of both video and natural language descriptions. We also introduce two tasks for video-and-language research based on VATEX: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2) Video-guided Machine Translation, to translate a source language description into the target language using the video information as additional spatiotemporal context. Extensive experiments on the VATEX dataset show that, first, the unified multilingual model can not only produce both English and Chinese descriptions for a video more efficiently, but also offer improved performance over the monolingual models. Furthermore, we demonstrate that the spatiotemporal video context can be effectively utilized to align source and target languages and thus assist machine translation. In the end, we discuss the potentials of using VATEX for other video-and-language research.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
ç§ãã¡ã¯ãè±èªãšäžåœèªã®äž¡æ¹ã§ 41,250 以äžã®ãããªãš 825,000 ã®ãã£ãã·ã§ã³ãå«ããæ°ããå€§èŠæš¡ãªå€èšèªãããªèšè¿°ããŒã¿ã»ãã VATEX ã玹ä»ããŸãããã£ãã·ã§ã³ã®äžã«ã¯ã206,000 以äžã®è±èªãšäžåœèªã®å¯Ÿèš³ãå«ãŸããŠããŸããåºã䜿çšãããŠãã MSR-VTT ããŒã¿ã»ãããšæ¯èŒããŠãVATEX ã¯å€èšèªå¯Ÿå¿ã§ãå€§èŠæš¡ã§ãèšèªçã«è€éã§ããããªãšèªç¶èšèªã®äž¡æ¹ã®èšè¿°ã®ç¹ã§ãã倿§ã§ãããŸããVATEX ã«åºã¥ããããªãšèšèªã®ç ç©¶ã®ããã® 2 ã€ã®ã¿ã¹ã¯ã玹ä»ããŸãã(1) ã³ã³ãã¯ããªçµ±äžãã£ãã·ã§ã³ ã¢ãã«ã䜿çšããŠããŸããŸãªèšèªã§ãããªãèšè¿°ããããšãç®çãšããå€èšèªãã㪠ãã£ãã·ã§ã³ãããã³ (2) ãããªã翻蚳ããããã®ãããªã¬ã€ãä»ãæ©æ¢°ç¿»èš³ãããªæ å ±ã远å ã®æç©ºéã³ã³ããã¹ããšããŠäœ¿çšããŠããœãŒã¹èšèªã®èª¬æãã¿ãŒã²ããèšèªã«å€æããŸããVATEX ããŒã¿ã»ããã«é¢ããåºç¯ãªå®éšã«ããããŸããçµ±åå€èšèªã¢ãã«ã¯ãããªã®è±èªãšäžåœèªã®äž¡æ¹ã®èª¬æãããå¹ççã«çæã§ããã ãã§ãªããåèšèªã¢ãã«ãããããã©ãŒãã³ã¹ãåäžããããšãããããŸãããããã«ãæç©ºéãããªã³ã³ããã¹ãã广çã«å©çšããŠãœãŒã¹èšèªãšã¿ãŒã²ããèšèªã調æŽããæ©æ¢°ç¿»èš³ãæ¯æŽã§ããããšãå®èšŒããŸããæåŸã«ãVATEX ãä»ã®ãããªãšèšèªã®ç ç©¶ã«äœ¿çšããå¯èœæ§ã«ã€ããŠèª¬æããŸãã
Perception Test MCQA
Perception Test MCQAã¯ãGoogle DeepMindã®Research Scientistã§ããViorica PÄtrÄuceanãã«ãã£ãŠã2023幎ã«èæ¡ããããã«ãã¢ãŒãã«ãããªã®ç¥èŠãšæšè«ã¹ãã«ãè©äŸ¡ããããã®ãã³ãããŒã¯ã§ããäžçäžã®çŽ100人ã®åå è ã«ãã£ãŠæ®åœ±ããããå¹³åé· 23ç§ã® 11.6kã®çŸå®äžçã®ãããªã䜿ãããŠããŸãã人éãšæå 端ã®ãã㪠QAã¢ãã«ãšã®ããã©ãŒãã³ã¹ã«ã¯å€§ããªå·®ããã(91.4%察46.2%)ããã«ãã¢ãŒãã«ãããªã®çè§£ã«ã¯å€§ããªæ¹åã®äœå°ãããããšã瀺åãããŠããŸãã
Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Abstract:
We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on computational tasks (e.g. classification, detection or tracking), the Perception Test focuses on skills (Memory, Abstraction, Physics, Semantics) and types of reasoning (descriptive, explanatory, predictive, counterfactual) across video, audio, and text modalities, to provide a comprehensive and efficient evaluation tool. The benchmark probes pre-trained models for their transfer capabilities, in a zero-shot / few-shot or limited finetuning regime. For these purposes, the Perception Test introduces 11.6k real-world videos, 23s average length, designed to show perceptually interesting situations, filmed by around 100 participants worldwide. The videos are densely annotated with six types of labels (multiple-choice and grounded video question-answers, object and point tracks, temporal action and sound segments), enabling both language and non-language evaluations. The fine-tuning and validation splits of the benchmark are publicly available (CC-BY license), in addition to a challenge server with a held-out test split. Human baseline results compared to state-of-the-art video QA models show a substantial gap in performance (91.4% vs 46.2%), suggesting that there is significant room for improvement in multimodal video understanding.
Dataset, baseline code, and challenge server are available at this https URLã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
æã ã¯ãäºåãã¬ãŒãã³ã°ããããã«ãã¢ãŒãã« ã¢ãã« (FlamingoãSeViLAãGPT-4 ãªã©) ã®ç¥èŠãšæšè«ã¹ãã«ãè©äŸ¡ããããã®ãæ°ãããã«ãã¢ãŒãã« ãã㪠ãã³ãããŒã¯ã§ããç¥èŠãã¹ããææ¡ããŸããèšç®ã¿ã¹ã¯ (åé¡ãæ€åºã远跡ãªã©) ã«çŠç¹ãåœãŠãæ¢åã®ãã³ãããŒã¯ãšæ¯èŒããŠãç¥èŠãã¹ãã¯ããããªããªãŒãã£ãªã«ãããã¹ãã« (èšæ¶ãæœè±¡åãç©çåŠãæå³è«) ãšæšè«ã®çš®é¡ (èšè¿°çã説æçãäºæž¬çãåäºå®ç) ã«çŠç¹ãåœãŠãŠããŸãã ãããã³ããã¹ã ã¢ããªãã£ã䜿çšããŠãå æ¬çã§å¹ççãªè©äŸ¡ããŒã«ãæäŸããŸãããã®ãã³ãããŒã¯ã¯ããŒãã·ã§ãã/å°æ°ã·ã§ããããŸãã¯éå®ããã埮調æŽäœå¶ã§ãäºåãã¬ãŒãã³ã°ãããã¢ãã«ã®è»¢éèœåã調æ»ããŸãããããã®ç®çã®ããã«ãç¥èŠãã¹ãã§ã¯ãäžçäžã®çŽ 100 人ã®åå è ã«ãã£ãŠæ®åœ±ããããç¥èŠçã«è峿·±ãç¶æ³ã瀺ãããã«èšèšããããå¹³åé· 23 ç§ã® 11.6k ã®çŸå®äžçã®ãããªãå°å ¥ãããŠããŸãããããªã«ã¯ 6 çš®é¡ã®ã©ãã« (å€è¢éžæåŒããã³æ ¹æ ã®ãããããªã®è³ªåãšåçããªããžã§ã¯ããšãã€ã³ãã®ãã©ãã¯ãäžæçãªã¢ã¯ã·ã§ã³ãšãµãŠã³ãã®ã»ã°ã¡ã³ã) ãå¯ã«æ³šéä»ããããŠãããèšèªãšéèšèªã®äž¡æ¹ã®è©äŸ¡ãå¯èœã§ãããã³ãããŒã¯ã®åŸ®èª¿æŽãšæ€èšŒã®åå²ã¯ãå ¬éããããã¹ãåå²ãåãããã£ã¬ã³ãž ãµãŒããŒã«å ããŠãå ¬éãããŠããŸã (CC-BY ã©ã€ã»ã³ã¹)ãæå 端ã®ãã㪠QA ã¢ãã«ãšæ¯èŒãã人éã®ããŒã¹ã©ã€ã³çµæã§ã¯ãããã©ãŒãã³ã¹ã«å€§ããªå·® (91.4% 察 46.2%) ã瀺ãããŠããããã«ãã¢ãŒãã« ãããªã®çè§£ã«ã¯å€§ããªæ¹åã®äœå°ãããããšã瀺åãããŠããŸãã
ããŒã¿ã»ãããããŒã¹ã©ã€ã³ ã³ãŒãããã£ã¬ã³ãž ãµãŒããŒã¯ããã® https URLããå ¥æã§ããŸãã
Audio
CoVoST2(21 languages)
CoVoST2ã¯Facebook AI Researchã®research engineerã®Changhan Wangãã«ãã£ãŠã2022幎ã«èæ¡ãããé³å£°ç¿»èš³ïŒSpeech translationãããã¯Speech-To-TextïŒã«é¢ãããã³ãããŒã¯ã§ããåŸæ¥ã®ããŒã¿ã»ãããšæ¯èŒããŠã倿°ã®èšèªã«å¯Ÿå¿ããŠããããšãç¹åŸŽã§ãã(21èšèªããè±èªãžã®ç¿»èš³ãããã³è±èªãã15èšèªãžã®ç¿»èš³)
CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Abstract:
Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in massive multilingual speech translation and speech translation for low resource language pairs, we release CoVoST 2, a large-scale multilingual speech translation corpus covering translations from 21 languages into English and from English into 15 languages. This represents the largest open dataset available to date from total volume and language coverage perspective. Data sanity checks provide evidence about the quality of the data, which is released under CC0 license. We also provide extensive speech recognition, bilingual and multilingual machine translation and speech translation baselines with open-source implementation.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
é³å£°ç¿»èš³ã¯ããã³ãããŒã¯ ããŒã¿ã»ããã®éçºã®åœ±é¿ããããæè¿ãŸããŸã人æ°ã®ããç ç©¶ããŒãã«ãªã£ãŠããŸããããã«ãããããããçŸåšã®ããŒã¿ã»ããã¯éãããæ°ã®èšèªãã«ããŒããŠããŸããå€§èŠæš¡ãªå€èšèªé³å£°ç¿»èš³ããã³ãªãœãŒã¹ã®å°ãªãèšèªãã¢ã®é³å£°ç¿»èš³ã®ç ç©¶ãä¿é²ããããšãç®çãšããŠã21 èšèªããè±èªãžã®ç¿»èš³ãããã³è±èªãã 15 èšèªãžã®ç¿»èš³ãã«ããŒããå€§èŠæš¡ãªå€èšèªé³å£°ç¿»èš³ã³ãŒãã¹ã§ãã CoVoST 2 ããªãªãŒã¹ããŸããããã¯ãç·éãšèšèªç¯å²ã®èгç¹ããããããŸã§ã«å©çšå¯èœãªæå€§ã®ãªãŒãã³ ããŒã¿ã»ããã«çžåœããŸããããŒã¿å¥å šæ§ãã§ãã¯ã¯ãCC0 ã©ã€ã»ã³ã¹ã«åºã¥ããŠãªãªãŒã¹ãããããŒã¿ã®å質ã«é¢ãã蚌æ ãæäŸããŸãããŸããåºç¯ãªé³å£°èªèãäºèšèªããã³å€èšèªã®æ©æ¢°ç¿»èš³ãããã³ãªãŒãã³ãœãŒã¹å®è£ ã«ããé³å£°ç¿»èš³ã®ããŒã¹ã©ã€ã³ãæäŸããŸãã
FLEURS(62 lang)
FLEURSã¯Meta AI Researchã®Alexis Conneauãã«ãã£ãŠã2023幎ã«èæ¡ãããé³å£°ã¿ã¹ã¯ã«é¢ãããã³ãããŒã¯ã§ããèªåé³å£°èªè (ASR)ãé³å£°èšèªèå¥ (Speech LangID)ãç¿»èš³ãæ€çŽ¢ãªã©ã®æ§ã ãªé³å£°ã¿ã¹ã¯ããããŸãã
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Abstract:
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic Speech Recognition (ASR), Speech Language Identification (Speech LangID), Translation and Retrieval. In this paper, we provide baselines for the tasks based on multilingual pre-trained models like mSLAM. The goal of FLEURS is to enable speech technology in more languages and catalyze research in low-resource speech understanding.ã¢ãã¹ãã©ã¯ã(æ©æ¢°ç¿»èš³)ïŒ
é³å£°ã®æ®éç衚çŸã®å°æ°ã·ã§ããåŠç¿è©äŸ¡ãã³ãããŒã¯ã§ãã FLEURS ã玹ä»ããŸããFLEURS ã¯ãæ©æ¢°ç¿»èš³ FLoRes-101 ãã³ãããŒã¯ãããŒã¹ã«æ§ç¯ããã 102 èšèªã® n-way 䞊åé³å£°ããŒã¿ã»ããã§ãèšèªããšã«çŽ 12 æéã®é³å£°ç£èŠãè¡ãããŸããFLEURS ã¯ãèªåé³å£°èªè (ASR)ãé³å£°èšèªèå¥ (Speech LangID)ãç¿»èš³ãæ€çŽ¢ãªã©ã®ããŸããŸãªé³å£°ã¿ã¹ã¯ã«äœ¿çšã§ããŸãããã®ããŒããŒã§ã¯ãmSLAM ã®ãããªäºåãã¬ãŒãã³ã°æžã¿ã®å€èšèªã¢ãã«ã«åºã¥ããŠã¿ã¹ã¯ã®ããŒã¹ã©ã€ã³ãæäŸããŸããFLEURS ã®ç®æšã¯ãããå€ãã®èšèªã§é³å£°ãã¯ãããžãŒãæå¹ã«ããäœãªãœãŒã¹ã®é³å£°çè§£ã®ç ç©¶ãä¿é²ããããšã§ãã
ãã³ãããŒã¯ã®æŠèŠããŸãšããŠã¿ãŠ
Geminiã®æ§èœè©äŸ¡ã«äœ¿ãããæ§ã
ãªãã³ãããŒã¯ããŸãšããŠã¿ãŠãæ¢åã®å€§èŠæš¡èšèªã¢ãã«ãå€§èŠæš¡ãã«ãã¢ãŒãã«ã¢ãã«ã人éãšæ¯èŒããŠãŸã ãŸã å£ã£ãŠããã¿ã¹ã¯ãæ°å€ãããããšãç¥ããŸããã
ãŸãããããã£ããã³ãããŒã¯ã®ç ç©¶ãšéçºããæ¢åã¢ãã«ã®åŒ±ç¹ãçºèŠã»ææããŠAGIã«è¿ã¥ãããã«å¿
èŠãªä»äºã§ããããšã宿ã§ããŸããã
ä»åŸãæ§ã ãªå€§èŠæš¡AIã¢ãã«ããªãªãŒã¹ãããŠãããšæããŸããããã®éã«ã¯ãã³ãããŒã¯ã®ã¹ã³ã¢ãã¿ãŠå·éã«æ¯èŒæ€èšãããŠãããããšæããŸãã


