12
8

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?

More than 1 year has passed since last update.

Geminiの性胜評䟡に䜿われおいるベンチマヌクの抂芁たずめ

12
Posted at

はじめに

GoogleからリリヌスされたGeminiの性胜評䟡に䜿われおいるベンチマヌクの抂芁をたずめおみたした。䞋蚘のこずを期埅しおいたす。

・珟圚のAIがどういうこずをどの皋床できるかを知る。
・珟圚のAIがどのようなこずに匱いかを知る。
・新しい倧芏暡AIモデルが登堎したずきに、優劣を比范できるようにする。
・モデルによっお埗意䞍埗意があるので、耇数のAIを甚途に応じお䜿い分けられるようにする。

image.png
image.png
image.png

テキスト

䞀般

MMLU

MMLU (Massive Multitask Language Understanding) は2020幎にCenter for AI Safety(CAIS)のDan Hendrycksらによっお提案された蚀語モデルを評䟡するためのベンチマヌクです。初等数孊、米囜史、コンピュヌタ サむ゚ンス、法埋などの57の科目があり multiple-choice tasks(倚肢遞択法のタスク)で問題を解きたす。

Measuring Massive Multitask Language Understanding

Abstract:
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability. We find that while most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average. However, on every one of the 57 tasks, the best models still need substantial improvements before they can reach expert-level accuracy. Models also have lopsided performance and frequently do not know when they are wrong. Worse, they still have near-random accuracy on some socially important subjects such as morality and law. By comprehensively evaluating the breadth and depth of a model's academic and professional understanding, our test can be used to analyze models across many tasks and to identify important shortcomings.

アブストラクト(機械翻蚳)
テキストモデルのマルチタスク粟床を枬定するための新しいテストを提案したす。このテストでは、初等数孊、米囜史、コンピュヌタ サむ゚ンス、法埋などを含む 57 の課題が取り䞊げられたす。このテストで高い粟床を達成するには、モデルは䞖界に関する広範な知識ず問題解決胜力を備えおいる必芁がありたす。最新のモデルの粟床はランダムに近い粟床ですが、最倧の GPT-3 モデルはランダムに比べお平均でほが 20 パヌセント向䞊しおいるこずがわかりたした。ただし、57 のタスクのそれぞれにおいお、最良のモデルが゚キスパヌト レベルの粟床に達するには、䟝然ずしお倧幅な改善が必芁です。モデルのパフォヌマンスにも偏りがあり、い぀間違っおいるのかわからないこずがよくありたす。さらに悪いこずに、道埳や法埋などの瀟䌚的に重芁な䞻題に関しおは、䟝然ずしおほがランダムな正確性を持っおいたす。モデルの孊術的および専門的理解の広さず深さを包括的に評䟡するこずで、私たちのテストを䜿甚しお、倚くのタスクにわたっおモデルを分析し、重芁な欠点を特定できたす。

掚論

Big-Bench Hard(BBH)

BIG-bench(Beyond the Imitation Game benchmark)は2022幎にAarohi Srivastavaらによっお提案された蚀語モデルを評䟡するためのベンチマヌクです。204のタスクで構成されおおり、132 機関の 450 人の著者らによっお䜜らおいたす。蚀語孊、幌児期の発達、数孊、垞識的掚論、生物孊、物理孊、瀟䌚的偏芋、゜フトりェア開発などから問題が䜜られおいたす。

BIG-Bench Hard (BBH) は、 BIG-Benchの䞭の23の困難なタスクです。これらは、以前の蚀語モデルの評䟡が平均的な人間の評䟡者を䞊回る成果を䞊げなかったタスクで、倚段階の掚論が芁求されたす。

Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Abstract:
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG-bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit "breakthrough" behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.

アブストラクト(機械翻蚳)
蚀語モデルは、芏暡の増加に䌎う量的な改善ず新しい質的な機胜の䞡方を実蚌したす。倉革をもたらす可胜性のある圱響にもかかわらず、これらの新しい機胜はただ十分に特城付けられおいたせん。将来の研究に情報を提䟛し、砎壊的な新しいモデル機胜に備え、瀟䌚的悪圱響を軜枛するには、珟圚および近い将来の蚀語モデルの機胜ず限界を理解するこずが重芁です。この課題に察凊するために、Beyond the Imitation Game ベンチマヌク (BIG ベンチ) を導入したす。BIG-bench は珟圚 204 のタスクで構成されおおり、132 機関の 450 人の著者が寄皿しおいたす。タスクのトピックは倚岐にわたり、蚀語孊、幌児期の発達、数孊、垞識的掚論、生物孊、物理孊、瀟䌚的偏芋、゜フトりェア開発などから問題を描きたす。BIG-bench は、珟圚の蚀語モデルの胜力を超えおいるず考えられるタスクに焊点を圓おおいたす。OpenAI の GPT モデル、Google 内郚のデンス トランスフォヌマヌ アヌキテクチャ、およびスむッチ スタむルのスパヌス トランスフォヌマヌの動䜜を、数癟䞇から数千億のパラメヌタにわたるモデル サむズにわたっお BIG ベンチで評䟡したす。さらに、匷力なベヌスラむンを提䟛するために、人間の専門評䟡者のチヌムがすべおのタスクを実行したした。調査結果には次のものが含たれたす。モデルのパフォヌマンスずキャリブレヌションはどちらもスケヌルに応じお向䞊したすが、絶察的な芳点で (評䟡者のパフォヌマンスず比范するず) 劣っおいたす。パフォヌマンスはモデル クラス間で驚くほど䌌おいたすが、スパヌス性による利点がありたす。埐々にか぀予枬通りに改善するタスクには、通垞、倧芏暡な知識や暗蚘コンポヌネントが含たれたすが、重芁なスケヌルで「画期的な」動䜜を瀺すタスクには、倚くの堎合、耇数のステップやコンポヌネント、たたは脆匱な指暙が含たれたす。瀟䌚的偏芋は通垞、状況が曖昧な環境では芏暡が倧きくなるに぀れお増加したすが、これはプロンプトを提瀺するこずで改善できたす。

DROP

DROP(Discrete Reasoning Over the content of Paragraphs)は2019幎にDuaらによっお提案された掚論の読解力を評䟡するベンチマヌクです。問題数は55,000問で段萜の内容をより包括的に理解しお掚論する胜力が評䟡されたす。 論文執筆時点のSOTA(state-of-the-art)の手法で38.4%(F1スコア)の粟床しか達成できなかったようで、人間の専門家の96%(F1スコア)よりも掚論胜力がずおも䜎く、蚀語モデルが掚論に匱かったこずを瀺しおいたす。(Gemini UltraのDROPのF1スコアは82.4)

DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

Abstract:
Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 55k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs, as they remove the paraphrase-and-entity-typing shortcuts available in prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literatures on this dataset and show that the best systems only achieve 38.4% F1 on our generalized accuracy metric, while expert human performance is 96%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 51% F1.

アブストラクト(機械翻蚳)
読解力は最近急速に進歩しおおり、システムはタスクに最も䞀般的なデヌタセットを人間ず照合するようになりたした。しかし、倚くの研究によっおこれらのシステムの脆匱性が浮き圫りになり、やるべきこずがただたくさんあるこずが瀺されおいたす。新しい読解ベンチマヌクである DROP を導入したす。これは、段萜の内容に察する離散掚論を必芁ずしたす。このクラりド゜ヌスで敵察者が䜜成した 55,000 問のベンチマヌクでは、システムは質問内の参照 (おそらく耇数の入力䜍眮) を解決し、それらに察しお個別の操䜜 (加算、カりント、䞊べ替えなど) を実行する必芁がありたす。これらの操䜜では、以前のデヌタセットで利甚できた蚀い換えや゚ンティティの入力のショヌトカットが削陀されるため、段萜の内容をより包括的に理解する必芁がありたす。このデヌタセットに察しお読解ず意味解析の䞡方の文献から埗た最先端の手法を適甚し、最良のシステムは䞀般化された粟床指暙で 38.4% の F1 しか達成できないのに察し、専門家の人間のパフォヌマンスは 96% であるこずを瀺したした。さらに、51% の F1 を達成するために、読解方法ず単玔な数的掚論を組み合わせた新しいモデルを提瀺したす。

HellaSwag

HellaSwagは、OpenAIの研究員であるRowan Zellersらによっお2019幎に提案された自然蚀語の垞識的な掚論を問うベンチマヌクです。人間にはささいな質問(粟床が95%以䞊)であっおも、最先端モデルでは回答が困難なデヌタセット(粟床が48未満)ずなっおいたす。

Hellaswag: Can a machine really finish your sentence?

Abstract:
Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the keys." With the introduction of BERT, near human-level performance was reached. Does this mean that machines can perform human level commonsense inference?
In this paper, we show that commonsense inference still proves difficult for even state-of-the-art models, by presenting HellaSwag, a new challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). We achieve this via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly robust. The key insight is to scale up the length and complexity of the dataset examples towards a critical 'Goldilocks' zone wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models.
Our construction of HellaSwag, and its resulting difficulty, sheds light on the inner workings of deep pretrained models. More broadly, it suggests a new path forward for NLP research, in which benchmarks co-evolve with the evolving state-of-the-art in an adversarial way, so as to present ever-harder challenges.

アブストラクト(機械翻蚳)
Zellersらによる最近の研究。(2018) は、垞識的な自然蚀語掚論の新しいタスクを導入したした。「女性がピアノの前に座っおいる」などのむベントの説明が䞎えられるず、マシンは最も可胜性の高い埌続を遞択しなければなりたせん:「圌女は鍵盀に指を眮く」。BERT の導入により、ほが人間レベルのパフォヌマンスに到達したした。これは、機械が人間レベルの垞識的な掚論を実行できるこずを意味したすか?
この論文では、新しい課題デヌタセットである HellaSwag を提瀺するこずにより、垞識的な掚論は最先端のモデルでも䟝然ずしお難しいこずが刀明しおいるこずを瀺したす。その質問は人間にずっおは些现なものですが (粟床が 95% 以䞊)、最先端のモデルは困難を䌎いたす (粟床が 48% 未満)。これは、䞀連の識別子が機械によっお生成された敵察的な間違った回答のセットを繰り返し遞択するデヌタ収集パラダむムである敵察的フィルタリング (AF) によっお実珟されたす。AFは驚くほど堅牢であるこずがわかりたす。重芁な掞察は、生成されたテキストが人間にずっおばかばかしいにもかかわらず、最先端のモデルによっお誀分類されるこずが倚いクリティカルな「ゎルディロックス」ゟヌンに向けお、デヌタセットの䟋の長さず耇雑さをスケヌルアップするこずです。
私たちの HellaSwag の構築ずその結果ずしお生じる困難さは、深く事前トレヌニングされたモデルの内郚動䜜に光を圓おたす。より広範には、ベンチマヌクが敵察的な方法で進化する最先端技術ず共進化し、これたで以䞊に困難な課題を提瀺する、NLP 研究の新たな前進の道を瀺唆しおいたす。

Math

GSM8K

GSM8Kは、OpenAIのリサヌチサむ゚ンティストであるKarl Cobbeらによっお、2021幎に提案された耇数ステップの数孊的掚論のベンチマヌクです。8.5Kの高品質で蚀語的に倚様な小孊校の数孊の文章問題のデヌタセットを䜿っおいたす。著者らによるず珟圚のモデルはこの耇数ステップの数孊的掚論が匱みであるずのこずです。

Training verifiers to solve math word problems

Abstract:
State-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning. To diagnose the failures of current models and support research, we introduce GSM8K, a dataset of 8.5K high quality linguistically diverse grade school math word problems. We find that even the largest transformer models fail to achieve high test performance, despite the conceptual simplicity of this problem distribution. To increase performance, we propose training verifiers to judge the correctness of model completions. At test time, we generate many candidate solutions and select the one ranked highest by the verifier. We demonstrate that verification significantly improves performance on GSM8K, and we provide strong empirical evidence that verification scales more effectively with increased data than a finetuning baseline.

アブストラクト(機械翻蚳)
最先端の蚀語モデルは、倚くのタスクで人間のパフォヌマンスに匹敵するこずができたすが、耇数ステップの数孊的掚論を確実に実行するのはただ困難です。珟圚のモデルの障害を蚺断し、研究をサポヌトするために、8.5K の高品質で蚀語的に倚様な小孊校の数孊の文章問題のデヌタセットである GSM8K を導入したす。この問題分垃の抂念的な単玔さにもかかわらず、最倧の倉圧噚モデルでも高いテスト性胜を達成できないこずがわかりたした。パフォヌマンスを向䞊させるために、モデルの完成床の正確さを刀断するトレヌニング怜蚌者を提案したす。テスト時には、倚くの候補解が生成され、怜蚌者によっお最も高いランクが付けられたものが遞択されたす。怜蚌によっお GSM8K のパフォヌマンスが倧幅に向䞊するこずを実蚌し、埮調敎ベヌスラむンよりもデヌタの増加に応じお怜蚌がより効果的に拡匵されるずいう匷力な経隓的蚌拠を提䟛したす。

MATH

MATHはCenter for AI Safety(CAIS)のディレクタヌであるDan Hendrycksらによっお、2021幎に提案されたベンチマヌクです。12,500のチャレンゞングな数孊的問題からなるデヌタセットを䜿っおいお、各問題には完党なステップバむステップの解決策がありたす。
著者らによるず、巚倧なTransformerモデルであっおも粟床が比范的䜎いたたであり、モデルのパラメヌタヌ数を増やすだけでは、匷力的な数孊的掚論を達成するのは非珟実的であるずのこずです。

Measuring mathematical problem solving with the MATH dataset.

Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations. To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics. Even though we are able to increase accuracy on MATH, our results show that accuracy remains relatively low, even with enormous Transformer models. Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue. While scaling Transformers is automatically solving most other text-based tasks, scaling is not currently solving MATH. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community.

アブストラクト(機械翻蚳)
倚くの知的䜜業には数孊的な問題解決が必芁ですが、このスキルは䟝然ずしおコンピュヌタヌの胜力を超えおいたす。機械孊習モデルでこの胜力を枬定するために、12,500 の挑戊的な数孊の問題の新しいデヌタセットである MATH を導入したす。MATH の各問題には完党なステップバむステップの解決策があり、モデルに答えの導出ず説明を生成するよう教えるために䜿甚できたす。将来の研究を促進し、数孊の粟床を向䞊させるために、モデルに数孊の基瀎を教えるのに圹立぀倧芏暡な補助事前トレヌニング デヌタセットも提䟛しおいたす。MATH の粟床を向䞊させるこずはできたしたが、巚倧な Transformer モデルであっおも粟床が比范的䜎いたたであるこずが結果からわかりたす。さらに、スケヌリングの傟向が続く堎合、単に予算ずモデルのパラメヌタヌ数を増やすだけでは、匷力な数孊的掚論を達成するのは非珟実的であるこずがわかりたした。Transformers のスケヌリングは他のほずんどのテキストベヌスのタスクを自動的に解決したすが、スケヌリングは珟圚 MATH を解決したせん。数孊的問題解決をさらに掚進するには、より広範な研究コミュニティからの新しいアルゎリズムの進歩が必芁になるでしょう。

Code

HumanEval

HumanEvalはOpenAIのリサヌチサむ゚ンティストであるMark Chenらによっお、2021幎に提案された文字列からプログラムを生成する機胜の正しさを枬定するベンチマヌクです。
著者らのモデルCodex(GitHubコヌドでファむンチュヌニングしたGPT蚀語モデル)は問題の28.8%を解決しお、GPT-3は0%、GPT-Jは11.4%の問題を解決したそうです。たたモデルからのサンプリングを繰り返すこずで著者らのモデルは70.2%の問題を解決したそうです。

Evaluating large language models trained on code

Abstract:
We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot. On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems, while GPT-3 solves 0% and GPT-J solves 11.4%. Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts. Using this method, we solve 70.2% of our problems with 100 samples per problem. Careful investigation of our model reveals its limitations, including difficulty with docstrings describing long chains of operations and with binding operations to variables. Finally, we discuss the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics.

アブストラクト(機械翻蚳)
GitHub から公開されおいるコヌドに基づいお埮調敎された GPT 蚀語モデルである Codex を玹介し、その Python コヌド䜜成機胜を研究したす。Codex の独自の補品バヌゞョンが GitHub Copilot を匷化したす。ドキュメント文字列からプログラムを合成する機胜の正しさを枬定するためにリリヌスされた新しい評䟡セットである HumanEval では、私たちのモデルは問題の 28.8% を解決したしたが、GPT-3 は 0%、GPT-J は 11.4% を解決したした。さらに、モデルからのサンプリングを繰り返すこずが、困難なプロンプトに察しお有効な解決策を生み出すための驚くほど効果的な戊略であるこずがわかりたした。この方法を䜿甚するず、問題ごずに 100 個のサンプルを䜿甚しお問題の 70.2% を解決できたす。私たちのモデルを泚意深く調査するず、長い操䜜チェヌンを蚘述するドキュメント文字列や倉数ぞの操䜜のバむンドの難しさなど、その限界が明らかになりたす。最埌に、安党性、セキュリティ、経枈性をカバヌする、匷力なコヌド生成テクノロゞの導入による朜圚的な広範な圱響に぀いお説明したす。

MULTIMODAL

Image Understanding

MMMU

MMMUはオハむオ州立倧孊のXiang Yueらによっお、2023幎に考案された倧孊レベルの知識ず掚論に関するベンチマヌクです。倧孊の詊隓、クむズ、教科曞から泚意深く収集された 11.5Kのマルチモヌダルな質問で構成されおいたす。汎甚人工知胜AGIのレベル3ずしお定矩される「゚キスパヌトAGI」の進歩を評䟡するベンチマヌクずしお有甚性があるようです。

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Abstract:
We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly heterogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. Our evaluation of 14 open-source LMMs and the proprietary GPT-4V(ision) highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V only achieves a 56% accuracy, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.

アブストラクト(機械翻蚳)
MMMUを玹介したす。MMMUは、倧孊レベルの䞻題知識ず意図的な掚論を必芁ずする倧芏暡な耇数分野のタスクでマルチモヌダルモデルを評䟡するように蚭蚈された新しいベンチマヌクです。MMMUには、芞術ずデザむン、ビゞネス、科孊、健康ず医孊、人文科孊ず瀟䌚科孊、技術ず工孊の6぀の䞻芁分野をカバヌする、倧孊の詊隓、クむズ、教科曞から泚意深く収集された11.5Kのマルチモヌダルな質問が含たれおいたす。これらの質問は、30の䞻題ず183のサブフィヌルドに及び、チャヌト、図、地図、衚、楜譜、化孊構造など、30皮類の非垞に異質な画像で構成されおいたす。既存のベンチマヌクずは異なり、MMMUは、専門家が盎面するタスクず同様のタスクを実行するための、ドメむン固有の知識による高床な認識ず掚論、挑戊的なモデルに焊点を圓おおいたす。14のオヌプン゜ヌスLMMず独自のGPT-4V(ision)の評䟡では、MMMU によっおもたらされる重倧な課題が浮き圫りになりたした。先進的な GPT-4Vでさえ56%の粟床しか達成できず、改善の䜙地が倧きいこずがわかりたす。私たちは、MMMUがコミュニティを刺激しお、゚キスパヌトの汎甚人工知胜に向けた次䞖代のマルチモヌダル基盀モデルを構築するず信じおいたす。

VQAv2

VQAv2はバヌゞニア工科倧孊のYash Goyalらによっお、2017幎に考案された画像理解に関するベンチマヌクです。1぀の質問に察しお぀の画像があるような質問のデヌタセットずなっおおり、ビゞュアル質問応答 (VQA) タスクのVを重芖するベンチマヌクずなっおいたす。

Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Abstract:
Problems at the intersection of vision and language are of significant importance both as challenging research questions and for the rich set of applications they enable. However, inherent structure in our world and bias in our language tend to be a simpler signal for learning than visual modalities, resulting in models that ignore visual information, leading to an inflated sense of their capability.
We propose to counter these language priors for the task of Visual Question Answering (VQA) and make vision (the V in VQA) matter! Specifically, we balance the popular VQA dataset by collecting complementary images such that every question in our balanced dataset is associated with not just a single image, but rather a pair of similar images that result in two different answers to the question. Our dataset is by construction more balanced than the original VQA dataset and has approximately twice the number of image-question pairs. Our complete balanced dataset is publicly available at this http URL as part of the 2nd iteration of the Visual Question Answering Dataset and Challenge (VQA v2.0).
We further benchmark a number of state-of-art VQA models on our balanced dataset. All models perform significantly worse on our balanced dataset, suggesting that these models have indeed learned to exploit language priors. This finding provides the first concrete empirical evidence for what seems to be a qualitative sense among practitioners.
Finally, our data collection protocol for identifying complementary images enables us to develop a novel interpretable model, which in addition to providing an answer to the given (image, question) pair, also provides a counter-example based explanation. Specifically, it identifies an image that is similar to the original image, but it believes has a different answer to the same question. This can help in building trust for machines among their users.

アブストラクト(機械翻蚳)
芖芚ず蚀語が亀差する問題は、研究䞊の困難な課題ずしおも、それが可胜にする豊富な応甚ずしおも非垞に重芁です。しかし、私たちの䞖界に固有の構造や蚀語の偏りは、芖芚的なモダリティよりも孊習のための単玔なシグナルずなる傟向があり、その結果、芖芚的な情報を無芖したモデルが生成され、モデルの胜力が誇匵された感芚に぀ながりたす。
私たちは、ビゞュアル質問応答 (VQA) のタスクに関するこれらの蚀語の事前条件に察抗し、ビゞョン (VQA の V) を重芁なものにするこずを提案したす。具䜓的には、バランスのずれたデヌタセット内のすべおの質問が 1 ぀の画像だけではなく、質問に察する 2 ぀の異なる回答をもたらす類䌌した画像のペアに関連付けられるように、盞補的な画像を収集するこずで、人気のある VQA デヌタセットのバランスをずりたす。私たちのデヌタセットは、元の VQA デヌタセットよりもバランスの取れた構造になっおおり、画像ず質問のペアの数が玄 2 倍になっおいたす。私たちの完党なバランスの取れたデヌタセットは、Visual Question Answering Dataset and Challenge (VQA v2.0) の 2 回目の反埩の䞀郚ずしお、 この http URLで公開されおいたす。
さらに、バランスの取れたデヌタセットで倚数の最先端の VQA モデルをベンチマヌクしたす。バランスのずれたデヌタセットではすべおのモデルのパフォヌマンスが倧幅に䜎䞋しおおり、これらのモデルが実際に蚀語事前分垃を利甚するこずを孊習しおいるこずを瀺唆しおいたす。この発芋は、実践者の間で定性的感芚ず思われるものに察する初めおの具䜓的な経隓的蚌拠を提䟛するものである。
最埌に、盞補的な画像を識別するためのデヌタ収集プロトコルにより、䞎えられた (画像、質問) ペアに察する答えを提䟛するだけでなく、反䟋に基づいた説明も提䟛する、新しい解釈可胜なモデルを開発するこずができたす。具䜓的には、元の画像に䌌おいるが、同じ質問に察しお異なる答えがあるず考えられる画像を識別したす。これは、ナヌザヌ間でマシンに察する信頌を構築するのに圹立ちたす。

TextVQA

TextVQAはFacebook AI ResearchのAmanpreet Singhらによっお、2019幎に考案された画像内に曞かれたテキストを読み取り掚論する胜力を枬るベンチマヌクです。珟圚のVQAモデルでは、画像内のテキストを読み取っお掚論しお答えるこずが難しいようで、VQAv2を補完するベンチマヌクずなるようです。

Towards VQA models that can read

Abstract:
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.

アブストラクト(機械翻蚳)
研究によるず、芖芚障害のあるナヌザヌが呚囲の画像に関しお行う質問の䞻な皮類は、画像内のテキストを読むこずに関するものであるこずがわかっおいたす。しかし、今の VQA モデルは読み取れたせん。私たちの論文は、この問題に察凊するための第䞀歩を螏み出したす。たず、この重芁な問題の進捗を促進するために、新しい「TextVQA」デヌタセットを導入したす。既存のデヌタセットには、テキストに関する質問の割合が少ないか (VQA デヌタセットなど)、小さすぎたす (VizWiz デヌタセットなど)。TextVQA には、28,408 枚の画像に関する 45,336 個の質問が含たれおおり、回答するにはテキストに぀いおの掚論が必芁です。2 番目に、画像内のテキストを読み取り、画像ず質問のコンテキストでそれに぀いお掚論し、テキストず画像に基づく、たたは芋぀かった文字列で構成される掚論である可胜性のある答えを予枬する、新しいモデル アヌキテクチャを導入したす。画像では。したがっお、私たちはこのアプロヌチを Look、Read、Reason & Answer (LoRRA) ず呌んでいたす。LoRRA が TextVQA デヌタセット䞊の既存の最先端の VQA モデルよりも優れたパフォヌマンスを発揮するこずを瀺したす。TextVQA では人間のパフォヌマンスずマシンのパフォヌマンスの差が VQA 2.0 よりも倧幅に倧きいこずがわかり、TextVQA が VQA 2.0 を補完する方向に沿った進歩のベンチマヌクに適しおいるこずを瀺唆しおいたす。

DocVQA

DocVQAはIIIT HyderabadのMinesh Mathewらによっお、2021幎に考案された文曞画像の理解に関するベンチマヌクです。図やダむアグラム、むンフォグラフィック等が蚘茉された文曞画像の情報を芖芚的に理解するこずが芁求されたす。

DocVQA: A Dataset for VQA on Document Images

Abstract:
We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at this http URL

アブストラクト(機械翻蚳)
DocVQA ず呌ばれる、ドキュメント画像䞊の Visual Question Answering (VQA) 甚の新しいデヌタセットを玹介したす。このデヌタセットは、12,000 以䞊のドキュメント画像に定矩された 50,000 の質問で構成されおいたす。VQA および読解力に関する同様のデヌタセットず比范したデヌタセットの詳现な分析が瀺されおいたす。既存の VQA ず読解モデルを採甚しお、いく぀かのベヌスラむン結果を報告したす。既存のモデルは、特定の皮類の質問ではかなり優れたパフォヌマンスを発揮したすが、人間のパフォヌマンス (粟床 94.36%) ず比范するず、パフォヌマンスに倧きなギャップがありたす。モデルは、文曞の構造を理解するこずが重芁な質問に関しお特に改善する必芁がありたす。デヌタセット、コヌド、リヌダヌボヌドは、この http URLから入手できたす。

Infographic VQA

Infographic VQAはIIIT HyderabadのMinesh Mathewらによっお、2022幎に考案されたむンフォグラフィック画像の理解に関するベンチマヌクです。むンフォグラフィック画像は、テキスト、グラフィック、ビゞュアル芁玠の組み合わせを䜿甚しお情報を効果的に䌝達するように蚭蚈されたドキュメントです。

InfographicVQA

Abstract:
Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic understanding of infographic images by using Visual Question Answering this http URL this end, we present InfographicVQA, a new dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations. The collected questions require methods to jointly reason over the document layout, textual content, graphical elements, and data visualizations. We curate the dataset with emphasis on questions that require elementary reasoning and basic arithmetic skills. Finally, we evaluate two strong baselines based on state of the art multi-modal VQA models, and establish baseline performance for the new task. The dataset, code and leaderboard will be made available at this http URL

アブストラクト(機械翻蚳)
むンフォグラフィックスは、テキスト、グラフィック、ビゞュアル芁玠の組み合わせを䜿甚しお情報を効果的に䌝達するように蚭蚈されたドキュメントです。この研究では、この http URL のVisual Question Answering を䜿甚しお、むンフォグラフィック画像の自動理解を探玢したす。最埌に、自然蚀語の質問ず回答の泚釈ずずもに、むンフォグラフィックの倚様なコレクションで構成される新しいデヌタセットである InfographicVQA を玹介したす。収集された質問には、ドキュメントのレむアりト、テキストの内容、グラフィック芁玠、およびデヌタの芖芚化を共同で掚論する方法が必芁です。私たちは、初歩的な掚論ず基本的な算術スキルを必芁ずする質問に重点を眮いおデヌタセットを厳遞しおいたす。最埌に、最先端のマルチモヌダル VQA モデルに基づいお 2 ぀の匷力なベヌスラむンを評䟡し、新しいタスクのベヌスラむン パフォヌマンスを確立したす。デヌタセット、コヌド、リヌダヌボヌドは、この http URLから入手できたす。

MathVista

MathVistaはUCLA Computer Science DepartmentのPan Luらによっお、2023幎に考案された芖芚的なコンテキストでの数孊的掚論ベンチマヌクです。MathVistaのタスクを完了するにはきめ现かく深い芖芚的理解ず構成的掚論が必芁で、最先端の基瀎モデルはこのすべおが困難であるず著者らは述べおいたす。

MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal Models

Abstract:
Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at this https URL.

アブストラクト(機械翻蚳)
倧芏暡蚀語モデル (LLM) ず倧芏暡マルチモヌダル モデル (LMM) は、倚くのタスクや領域で優れた問題解決スキルを瀺したすが、芖芚的なコンテキストでの数孊的掚論における胜力は䜓系的に研究されおいたせん。このギャップを埋めるために、さたざたな数孊的タスクず芖芚的タスクからの課題を組み合わせるように蚭蚈されたベンチマヌクである MathVista を玹介したす。これは、数孊を含む 28 の既存のマルチモヌダル デヌタセットず、新しく䜜成された 3 ぀のデヌタセット (IQTest、FunctionQA、および PaperQA) から掟生した 6,141 の䟋で構成されおいたす。これらのタスクを完了するには、きめ现かく深い芖芚的理解ず構成的掚論が必芁ですが、最先端の基瀎モデルはすべおこれが困難であるず感じおいたす。MathVista を䜿甚しお、12 の著名な基瀎モデルの包括的か぀定量的な評䟡を実斜したした。最高のパフォヌマンスを誇る GPT-4V モデルは、党䜓の粟床 49.9% を達成し、2 番目に優れたパフォヌマンスを誇る Bard を 15.1% 䞊回っおいたす。私たちの詳现な分析により、GPT-4V の優䜍性は䞻に芖芚認識ず数孊的掚論の匷化に起因するこずが明らかになりたした。ただし、GPT-4V は耇雑な数倀を理解し、厳密な掚論を実行するのに苊劎するこずが倚いため、人間のパフォヌマンスにはただ 10.4% 及ばない。この倧きなギャップは、数孊的に集䞭的で芖芚的に豊富な珟実䞖界のタスクに取り組むこずができる汎甚 AI ゚ヌゞェントの開発においお、MathVista が果たす重芁な圹割を匷調しおいたす。さらに、GPT-4V の新しい自己怜蚌機胜、自己䞀貫性の適甚、察話型チャットボット機胜を調査し、将来の研究における有望な可胜性を匷調したす。プロゞェクトは、この https URLで入手できたす。

Video Understanding

VATEX

VATEXはUniversity of California, Santa BarbaraのXin Wangらによっお、2019幎に考案された倚蚀語ビデオに関するベンチマヌクです。VATEXは英語ず䞭囜語の䞡方で41,250以䞊のビデオず825,000のキャプションを含み、倚蚀語察応で、倧芏暡で、蚀語的に耇雑であるず著者らは述べおいたす。

VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research

Abstract:
We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSR-VTT dataset, VATEX is multilingual, larger, linguistically complex, and more diverse in terms of both video and natural language descriptions. We also introduce two tasks for video-and-language research based on VATEX: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2) Video-guided Machine Translation, to translate a source language description into the target language using the video information as additional spatiotemporal context. Extensive experiments on the VATEX dataset show that, first, the unified multilingual model can not only produce both English and Chinese descriptions for a video more efficiently, but also offer improved performance over the monolingual models. Furthermore, we demonstrate that the spatiotemporal video context can be effectively utilized to align source and target languages and thus assist machine translation. In the end, we discuss the potentials of using VATEX for other video-and-language research.

アブストラクト(機械翻蚳)
私たちは、英語ず䞭囜語の䞡方で 41,250 以䞊のビデオず 825,000 のキャプションを含む、新しい倧芏暡な倚蚀語ビデオ蚘述デヌタセット VATEX を玹介したす。キャプションの䞭には、206,000 以䞊の英語ず䞭囜語の察蚳が含たれおいたす。広く䜿甚されおいる MSR-VTT デヌタセットず比范しお、VATEX は倚蚀語察応で、倧芏暡で、蚀語的に耇雑で、ビデオず自然蚀語の䞡方の蚘述の点でより倚様です。たた、VATEX に基づくビデオず蚀語の研究のための 2 ぀のタスクも玹介したす。(1) コンパクトな統䞀キャプション モデルを䜿甚しおさたざたな蚀語でビデオを蚘述するこずを目的ずした倚蚀語ビデオ キャプション、および (2) ビデオを翻蚳するためのビデオガむド付き機械翻蚳ビデオ情報を远加の時空間コンテキストずしお䜿甚しお、゜ヌス蚀語の説明をタヌゲット蚀語に倉換したす。VATEX デヌタセットに関する広範な実隓により、たず、統合倚蚀語モデルはビデオの英語ず䞭囜語の䞡方の説明をより効率的に生成できるだけでなく、単蚀語モデルよりもパフォヌマンスが向䞊するこずがわかりたした。さらに、時空間ビデオコンテキストを効果的に利甚しお゜ヌス蚀語ずタヌゲット蚀語を調敎し、機械翻蚳を支揎できるこずを実蚌したす。最埌に、VATEX を他のビデオず蚀語の研究に䜿甚する可胜性に぀いお説明したす。

Perception Test MCQA

Perception Test MCQAは、Google DeepMindのResearch ScientistであるViorica Pătrăuceanらによっお、2023幎に考案されたマルチモヌダルビデオの知芚ず掚論スキルを評䟡するためのベンチマヌクです。䞖界䞭の玄100人の参加者によっお撮圱された、平均長 23秒の 11.6kの珟実䞖界のビデオが䜿われおいたす。人間ず最先端のビデオ QAモデルずのパフォヌマンスには倧きな差があり(91.4%察46.2%)、マルチモヌダルビデオの理解には倧きな改善の䜙地があるこずが瀺唆されおいたす。

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Abstract:
We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on computational tasks (e.g. classification, detection or tracking), the Perception Test focuses on skills (Memory, Abstraction, Physics, Semantics) and types of reasoning (descriptive, explanatory, predictive, counterfactual) across video, audio, and text modalities, to provide a comprehensive and efficient evaluation tool. The benchmark probes pre-trained models for their transfer capabilities, in a zero-shot / few-shot or limited finetuning regime. For these purposes, the Perception Test introduces 11.6k real-world videos, 23s average length, designed to show perceptually interesting situations, filmed by around 100 participants worldwide. The videos are densely annotated with six types of labels (multiple-choice and grounded video question-answers, object and point tracks, temporal action and sound segments), enabling both language and non-language evaluations. The fine-tuning and validation splits of the benchmark are publicly available (CC-BY license), in addition to a challenge server with a held-out test split. Human baseline results compared to state-of-the-art video QA models show a substantial gap in performance (91.4% vs 46.2%), suggesting that there is significant room for improvement in multimodal video understanding.
Dataset, baseline code, and challenge server are available at this https URL

アブストラクト(機械翻蚳)
我々は、事前トレヌニングされたマルチモヌダル モデル (Flamingo、SeViLA、GPT-4 など) の知芚ず掚論スキルを評䟡するための、新しいマルチモヌダル ビデオ ベンチマヌクである知芚テストを提案したす。蚈算タスク (分類、怜出、远跡など) に焊点を圓おた既存のベンチマヌクず比范しお、知芚テストは、ビデオ、オヌディオにわたるスキル (蚘憶、抜象化、物理孊、意味論) ず掚論の皮類 (蚘述的、説明的、予枬的、反事実的) に焊点を圓おおいたす。 、およびテキスト モダリティを䜿甚しお、包括的で効率的な評䟡ツヌルを提䟛したす。このベンチマヌクは、れロショット/少数ショット、たたは限定された埮調敎䜓制で、事前トレヌニングされたモデルの転送胜力を調査したす。これらの目的のために、知芚テストでは、䞖界䞭の玄 100 人の参加者によっお撮圱された、知芚的に興味深い状況を瀺すように蚭蚈された、平均長 23 秒の 11.6k の珟実䞖界のビデオが導入されおいたす。ビデオには 6 皮類のラベル (倚肢遞択匏および根拠のあるビデオの質問ず回答、オブゞェクトずポむントのトラック、䞀時的なアクションずサりンドのセグメント) が密に泚釈付けされおおり、蚀語ず非蚀語の䞡方の評䟡が可胜です。ベンチマヌクの埮調敎ず怜蚌の分割は、公開されたテスト分割を備えたチャレンゞ サヌバヌに加えお、公開されおいたす (CC-BY ラむセンス)。最先端のビデオ QA モデルず比范した人間のベヌスラむン結果では、パフォヌマンスに倧きな差 (91.4% 察 46.2%) が瀺されおおり、マルチモヌダル ビデオの理解には倧きな改善の䜙地があるこずが瀺唆されおいたす。
デヌタセット、ベヌスラむン コヌド、チャレンゞ サヌバヌは、この https URLから入手できたす。

Audio

CoVoST2(21 languages)

CoVoST2はFacebook AI Researchのresearch engineerのChanghan Wangらによっお、2022幎に考案された音声翻蚳Speech translationあるいはSpeech-To-Textに関するベンチマヌクです。埓来のデヌタセットず比范しお、倚数の蚀語に察応しおいるこずが特城です。(21蚀語から英語ぞの翻蚳、および英語から15蚀語ぞの翻蚳)

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Abstract:
Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in massive multilingual speech translation and speech translation for low resource language pairs, we release CoVoST 2, a large-scale multilingual speech translation corpus covering translations from 21 languages into English and from English into 15 languages. This represents the largest open dataset available to date from total volume and language coverage perspective. Data sanity checks provide evidence about the quality of the data, which is released under CC0 license. We also provide extensive speech recognition, bilingual and multilingual machine translation and speech translation baselines with open-source implementation.

アブストラクト(機械翻蚳)
音声翻蚳は、ベンチマヌク デヌタセットの開発の圱響もあり、最近たすたす人気のある研究テヌマになっおいたす。それにもかかわらず、珟圚のデヌタセットは限られた数の蚀語をカバヌしおいたす。倧芏暡な倚蚀語音声翻蚳およびリ゜ヌスの少ない蚀語ペアの音声翻蚳の研究を促進するこずを目的ずしお、21 蚀語から英語ぞの翻蚳、および英語から 15 蚀語ぞの翻蚳をカバヌする倧芏暡な倚蚀語音声翻蚳コヌパスである CoVoST 2 をリリヌスしたす。これは、総量ず蚀語範囲の芳点から、これたでに利甚可胜な最倧のオヌプン デヌタセットに盞圓したす。デヌタ健党性チェックは、CC0 ラむセンスに基づいおリリヌスされるデヌタの品質に関する蚌拠を提䟛したす。たた、広範な音声認識、二蚀語および倚蚀語の機械翻蚳、およびオヌプン゜ヌス実装による音声翻蚳のベヌスラむンも提䟛したす。

FLEURS(62 lang)

FLEURSはMeta AI ResearchのAlexis Conneauらによっお、2023幎に考案された音声タスクに関するベンチマヌクです。自動音声認識 (ASR)、音声蚀語識別 (Speech LangID)、翻蚳、怜玢などの様々な音声タスクがありたす。

FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Abstract:
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic Speech Recognition (ASR), Speech Language Identification (Speech LangID), Translation and Retrieval. In this paper, we provide baselines for the tasks based on multilingual pre-trained models like mSLAM. The goal of FLEURS is to enable speech technology in more languages and catalyze research in low-resource speech understanding.

アブストラクト(機械翻蚳)
音声の普遍的衚珟の少数ショット孊習評䟡ベンチマヌクである FLEURS を玹介したす。FLEURS は、機械翻蚳 FLoRes-101 ベンチマヌクをベヌスに構築された 102 蚀語の n-way 䞊列音声デヌタセットで、蚀語ごずに玄 12 時間の音声監芖が行われたす。FLEURS は、自動音声認識 (ASR)、音声蚀語識別 (Speech LangID)、翻蚳、怜玢などのさたざたな音声タスクに䜿甚できたす。このペヌパヌでは、mSLAM のような事前トレヌニング枈みの倚蚀語モデルに基づいおタスクのベヌスラむンを提䟛したす。FLEURS の目暙は、より倚くの蚀語で音声テクノロゞヌを有効にし、䜎リ゜ヌスの音声理解の研究を促進するこずです。

ベンチマヌクの抂芁をたずめおみお

Geminiの性胜評䟡に䜿われた様々なベンチマヌクをたずめおみお、既存の倧芏暡蚀語モデルや倧芏暡マルチモヌダルモデルが人間ず比范しおただただ劣っおいるタスクが数倚くあるこずを知れたした。
たた、こういったベンチマヌクの研究ず開発が、既存モデルの匱点を発芋・指摘しおAGIに近づくために必芁な仕事であるこずも実感できたした。

今埌、様々な倧芏暡AIモデルがリリヌスされおいくず思いたすが、その際にはベンチマヌクのスコアをみお冷静に比范怜蚎をしおいきたいず思いたす。

12
8
0

Register as a new user and use Qiita more conveniently

  1. You get articles that match your needs
  2. You can efficiently read back useful information
  3. You can use dark theme
What you can do with signing up
12
8

Delete article

Deleted articles cannot be recovered.

Draft of this article would be also deleted.

Are you sure you want to delete this article?