メインコンテンツへスキップ
見出し画像

【aibuilderclub.com】AIはもう「プロンプト力」だけじゃない——自走するAIを設計する4段階

    AI活用で差がつくのは、うまい指示文を書ける人だけではありません。AIが正しい情報を受け取り、途中で迷わず、結果を検証されながら自動で動き続ける仕組みを作れる人です。Prompt→Context→Harness→Loopという4段階から、2026年のAI開発と仕事の未来を読み解きます。


    Owner: aibuilderclub.com
    タイトル: Prompt vs Context vs Harness vs Loop Engineering: The 4 Shifts
    日時: 2026/07/02(更新日/初出 2026/06/11)
    URL: “https://www.aibuilderclub.com/blog/prompt-context-harness-evolution”

    記事はAI Builder Clubの「Build AI Agents」コース内に掲載され、Shirley名義で公開されています。AI Builder Clubは、AIエージェント、Claude Code、MCP、周辺の実行基盤を扱う学習コミュニティ兼メディアです。


    🧩 英語要約

    The article explains the evolution of AI engineering through four nested layers: prompt engineering, context engineering, harness engineering, and loop engineering. Each layer emerged because increasingly complex tasks exposed the limits of the previous one.

    Prompt engineering focuses on expressing instructions clearly. Roles, examples, constraints, and output formats can significantly influence an LLM’s response. However, even a perfect prompt cannot supply information the model has never received. This limitation led to context engineering, which manages the documents, tools, memories, search results, and other information available to the model at each step.

    As AI systems began performing long sequences of tool calls, correct information alone was no longer enough. Agents could drift from the original goal, misunderstand tool results, repeat mistakes, or falsely claim completion. Harness engineering addresses these problems by designing the execution environment around the model. A harness may include tools, file systems, sandboxes, memory, orchestration logic, tests, evaluators, recovery procedures, and access controls.

    Loop engineering adds another layer above the harness. Instead of a human manually starting every task, a loop detects work, triggers an agent, checks the result, stores progress, and decides whether to retry, continue, escalate, or stop. Its essential components include a measurable goal, a verifier, persistent state, triggers, termination conditions, and cost controls.

    The central message is that these disciplines do not replace one another. A prompt exists inside a context; context management operates inside a harness; and the harness runs inside a loop. Reliable AI systems therefore depend less on a single clever instruction and more on the complete system surrounding the model.


    🧩 日本語解説

    この記事は、AI開発の中心が「プロンプトエンジニアリング」から「コンテキストエンジニアリング」「ハーネスエンジニアリング」、さらに「ループエンジニアリング」へと広がってきた過程を、4つの入れ子構造として説明しています。

    第1段階のプロンプトエンジニアリングは、AIへの依頼を分かりやすく表現する技術です。役割、例、制約、出力形式などを適切に指定すれば、同じモデルでも回答品質は大きく変わります。しかし、どれほど優れた指示を書いても、社内資料や最新情報など、AIに渡されていない事実を生み出すことはできません。

    そこで登場するのがコンテキストエンジニアリングです。これは、プロンプトだけでなく、会話履歴、検索結果、文書、ツールの説明、メモリなど、AIがその瞬間に参照できる情報全体を設計する考え方です。情報を大量に入れればよいわけではなく、注意力を浪費しないよう、必要な情報だけを適切なタイミングで渡すことが重要です。Anthropicも、コンテキストを有限な資源として扱い、高い信号を持つ最小限の情報を選ぶ必要があると説明しています。

    しかし、正しい情報を渡しても、長時間動くAIエージェントは途中で目的を見失ったり、ツールの結果を誤読したり、失敗を重ねたりします。ここで必要になるのがハーネスエンジニアリングです。ハーネスとは、モデルの外側にあるファイルシステム、実行ツール、サンドボックス、メモリ、権限管理、評価、テスト、エラー回復、複数エージェントの調整などを含む実行基盤です。LangChainはこれを「Agent=Model+Harness」と整理しています。

    さらに、ハーネスが一回の実行を安全に管理するのに対し、ループエンジニアリングは「いつ実行し、結果をどう判定し、次に何をするか」まで自動化します。スケジュールやイベントをきっかけに仕事を発見し、AIへ渡し、検証し、記録し、必要なら再実行します。人間が毎回スタートボタンを押さなくても、仕事が継続する仕組みです。

    ただし、自律ループには暴走コストや誤りの大量生産というリスクがあります。明確な終了条件、予算上限、検証器、失敗時の停止・人間への引き継ぎが欠かせません。記事の結論は、優れたAI開発とは「魔法のプロンプト」を探すことではなく、モデルを囲む情報・実行・検証・反復のシステム全体を設計することだ、という点にあります。


    🧩 CEFR B1以上の重要語彙 12語

    🧩 01|phrasing

    日本語訳:言い回し、表現方法
    Example: Changing the phrasing of the prompt produced a clearer answer.

    🧩 02|context

    日本語訳:文脈、背景情報
    Example: The agent needs enough context to understand the customer’s request.

    🧩 03|finite

    日本語訳:有限の、限りがある
    Example: An AI model has a finite amount of attention available for each task.

    🧩 04|curate

    日本語訳:情報を選び、整理する
    Example: Engineers must curate the information placed in the context window.

    🧩 05|harness

    日本語訳:制御する仕組み、実行基盤
    Example: The harness gives the model tools, memory, tests, and safety controls.

    🧩 06|persistent

    日本語訳:持続する、永続的な
    Example: Persistent memory allows an agent to continue its work across sessions.

    🧩 07|verify

    日本語訳:検証する、正しいか確かめる
    Example: The system must verify the code before deploying it to production.

    🧩 08|autonomous

    日本語訳:自律的な、自動で行動する
    Example: An autonomous agent can select and use tools without constant guidance.

    🧩 09|trigger

    日本語訳:きっかけ、作動させるもの
    Example: A new support ticket can trigger the agent’s investigation loop.

    🧩 10|termination

    日本語訳:終了、停止
    Example: Every autonomous loop needs a clear termination condition.

    🧩 11|constraint

    日本語訳:制約、制限条件
    Example: A spending limit is an important constraint for an AI agent.

    🧩 12|bottleneck

    日本語訳:進行を妨げる要因、ボトルネック
    Example: Human approval can become a bottleneck in a recurring workflow.


    🧩 外部参照情報

    🧩 01|YouTube関連動画

    Effective Context Engineering for AI Agents — Why Agents Still Fail in Practice

    AIエージェントに大量の情報を与えるだけでは不十分であり、必要な情報を適切なタイミングで管理する重要性を解説した関連動画です。

    🧩 02|Reddit(USA)関連トピック

    Is harness a new buzzword?

    「ハーネス」という言葉は単なる流行語なのか、それともモデルを実用化する周辺コードを表す有効な概念なのか、開発者が議論しています。


    🧩 地名・人名・キーワード調査

    🧩 記事内の地名

    明示的な地名はほぼなし
    この記事は地域や都市を扱う記事ではなく、AI開発手法の変化を説明する技術記事です。登場する企業や開発者の多くは米国を中心とする英語圏のAI・ソフトウェア開発コミュニティと関係しています。

    米国(関連する産業背景)
    Anthropic、OpenAI、LangChain、Googleなど、記事の議論を支える企業・プロジェクトの多くが米国のAI産業と深く結び付いています。そのため、記事の内容は米国のAIエージェント開発競争を強く反映していると考えられます。これは記事と参照先から導ける背景上の整理です。

    🧩 記事に登場する人物名

    Peter Steinberger
    オープンソースAIエージェント「OpenClaw」の開発者。AIへ一回ずつ指示を出すのではなく、複数のエージェントを動かすループそのものを設計する考え方を発信しました。2026年2月にはOpenAIへの参加を表明し、OpenClawは独立した財団の下でオープンなプロジェクトとして継続すると説明しています。

    Boris Cherny
    AnthropicのAIコーディングツール「Claude Code」の作者として知られる技術者です。記事では、個別のプロンプトを書くことよりも、継続的に動くループを設計することが仕事になりつつあると説明した人物として紹介されています。

    Addy Osmani
    長年GoogleでChromeの開発者体験やAI関連の技術普及を担当してきたエンジニアリング・DevRelリーダーです。「Loop Engineering」の記事で、自動実行、分離された作業環境、スキル、外部接続、サブエージェント、永続メモリなどの構成要素を整理しました。

    🧩 その他の主要キーワード

    RAG(Retrieval-Augmented Generation)
    AIが回答する前に、外部のデータベースや文書から関連情報を検索し、その情報をコンテキストへ追加する方式です。モデルを再学習しなくても、社内情報や最新情報を利用できる点が強みです。ただし、検索結果の選び方、文書の分割方法、順位付けが不適切だと、誤った情報や不要な情報が入り込みます。

    Verifier(検証器)
    AIの出力が「十分に正しいか」「目標を達成したか」を判定する仕組みです。コードであればテスト、文章であれば必要項目や引用元の確認、データ処理であれば計算結果の照合などが該当します。自律ループでは、AI本体よりも検証器の品質が信頼性を左右する場合があります。


    🧩 日本・米国の比較情報

    🧩 01|Japan / 日本

    English:

    AI adoption is steadily expanding in Japan, particularly among large enterprises. However, the 2026 survey released by Japan’s Information-technology Promotion Agency found that most reported benefits remain concentrated in operational efficiency and faster processes. Many companies recognize positive effects, but the proportion achieving results that meet or exceed their original expectations is still limited.

    This situation closely matches the article’s framework. Japanese organizations may already be experimenting with prompts and individual AI tools, but the next challenge is to redesign complete business processes. That requires reliable context pipelines, controlled execution environments, verification, persistent records, human escalation rules, and measurable business goals.

    The opportunity for Japan is not simply to introduce more chatbots. It is to move from isolated productivity improvements toward systems that create new services and business value. At the same time, strong governance, information security, cost controls, and clear responsibility will be essential when agents begin acting across internal systems. The IPA survey covered 1,799 Japanese companies between April and June 2026.

    日本語:

    日本では、特に大企業を中心にAI導入が着実に広がっています。ただし、IPAが2026年7月に公表した国内企業1,799社の調査では、AIの利用目的と効果は業務の効率化・迅速化に集中しており、期待どおり、または期待以上の効果を得た企業の割合はまだ限定的だとされています。

    これは記事の4段階モデルで考えると、日本企業の多くがプロンプトや個別ツールの利用から、業務全体の設計へ移る途中にあることを示します。今後は、社内情報を適切に渡すコンテキスト設計、権限や実行環境を管理するハーネス、成果を測定して改善を続けるループが重要になります。

    単純な文章作成の短縮だけでなく、顧客対応、品質管理、開発、営業支援などの業務プロセス全体を見直し、新しい価値へつなげられるかが課題です。同時に、機密情報、誤作動、責任分担、利用コストへの対策も欠かせません。

    🧩 02|United States / 米国

    English:

    AI use is more broadly visible across the US business landscape, although adoption varies significantly by company size and industry. US Census Bureau data collected from December 2025 through May 2026 showed that overall business AI use remained between 17% and 20%. Among firms with at least 250 employees, 37% reported using AI.

    Adoption was especially high in information businesses, at 39.7%, and finance and insurance, at 33.9% as of May 3, 2026. These figures cover AI use in any business function, not only autonomous agents, so they should not be interpreted as direct measures of harness or loop engineering.

    Nevertheless, the United States has a strong ecosystem of AI model companies, agent platforms, developer tools, cloud infrastructure, and startups. This gives businesses more opportunities to experiment with long-running agents and automated workflows. The challenge is shifting from experimentation to dependable operations. Security isolation, evaluation, audit logs, budget limits, and human escalation become increasingly important when an agent can modify files, execute code, contact services, or spend money without constant supervision.

    日本語:

    米国勢調査局の調査では、2025年12月から2026年5月までの企業におけるAI利用率は全体で17~20%でした。従業員規模が250人以上の企業では37%がAIを利用しており、2026年5月3日時点では情報産業が39.7%、金融・保険が33.9%となっています。

    ただし、この数字は「何らかの業務機能でAIを使っている企業」の割合であり、自律型AIエージェントやループエンジニアリングだけを測定した数字ではありません。

    米国には、基盤モデル、クラウド、AIエージェント、開発ツール、評価基盤を提供する企業が集まり、長時間動くエージェントや複数エージェントの実験を進めやすい環境があります。一方、AIにコード実行、ファイル変更、外部サービスへの接続を許可するほど、セキュリティ、監査記録、予算上限、停止条件、人間への引き継ぎが重要になります。


    🧩 応用・ディスカッション展開

    🧩 01|Theme 1: Which Tasks Should Companies Automate With Agent Loops?

    Companies should not begin by placing every recurring task inside an autonomous agent loop. The best starting points are tasks that are frequent, measurable, reversible, and supported by reliable data. Examples include sorting support tickets, checking whether websites contain broken links, monitoring test failures, preparing routine reports, or identifying documents that require human review.

    A task should have a clear definition of success. For example, “improve customer service” is too vague for a loop. A better goal would be, “classify new support requests within five minutes, attach the relevant account information, and send uncertain cases to a human employee.” This goal can be measured, tested, and safely interrupted.

    Companies should also consider the cost of mistakes. An agent that drafts an internal summary creates less risk than one that issues refunds, changes production databases, or contacts customers. High-risk actions should require stronger verification or human approval. The loop may still perform research and prepare recommendations, while a person remains responsible for the final action.

    Another important question is reversibility. Can the organization undo the agent’s action? File changes can be managed with version control, while financial transfers or public messages may be difficult to reverse. Logs and persistent records are necessary so that people can understand why the system acted.

    Therefore, the goal is not maximum autonomy. The goal is appropriate autonomy. A well-designed loop removes repetitive coordination while preserving human judgment at the points where values, uncertainty, and serious consequences are involved.

    🧩 02|Theme 2: Who Is Responsible When an Autonomous Agent Fails?

    Responsibility for an autonomous agent should remain with the organization and people who design, approve, deploy, and operate the system. Describing an AI agent as autonomous does not make it a responsible legal or moral actor. The agent follows a combination of model behavior, supplied information, tools, permissions, evaluation rules, and organizational decisions.

    Responsibility should therefore be divided clearly. Product owners define the business objective and acceptable risk. Engineers build the harness, access controls, recovery mechanisms, and monitoring. Domain specialists define what a correct result looks like. Security teams review data access and external connections. Managers decide where human approval is mandatory. Senior leadership remains responsible for the organization’s overall use of the system.

    A strong audit trail is essential. The system should record the goal, information supplied to the model, tool calls, outputs, verifier decisions, costs, failures, retries, and human interventions. Without these records, an organization cannot investigate an error or improve the process.

    The system must also know when to stop. Repeated failure, unexpected spending, conflicting instructions, missing data, or low verifier confidence should trigger an escalation to a person. A loop that continues operating after it loses reliable feedback is not autonomous in a useful sense; it is simply uncontrolled.

    Human oversight should not mean manually approving every harmless action. That would remove most of the benefit. Instead, oversight should be risk-based. Low-impact, reversible work can run automatically, while high-impact decisions require stronger checks. The final principle is simple: organizations may delegate work to AI, but they cannot delegate accountability.

    🧩 03|Theme 3: Will Prompt Engineering Become Obsolete?

    Prompt engineering will not disappear, but it will become one part of a larger discipline. A clear instruction still matters because even the best context pipeline, harness, or autonomous loop can fail when the goal is ambiguous. However, complex AI systems cannot be made reliable through wording alone.

    In a simple chat interaction, the user writes a prompt and immediately evaluates the answer. In an agent system, the model may call tools, read files, generate code, update records, and continue working for many steps. During that process, success depends on more than the original message. The system must decide what information to provide, which tools are available, how permissions work, what state persists, how results are checked, and when execution should stop.

    Prompt engineering is therefore becoming similar to interface design. Engineers must still describe goals and expected behavior, but they should not try to encode an entire application as one enormous instruction. Excessive rules can make prompts fragile and difficult to maintain. It is often better to place deterministic requirements in code, permissions, schemas, tests, and validators.

    Future AI professionals will need several connected skills: clear communication, information architecture, software engineering, evaluation design, security, workflow analysis, and cost management. Nontechnical employees will also benefit from understanding the difference between requesting an answer and designing a dependable process.

    Prompting remains the first floor of the building. Context, harnesses, and loops add higher floors, but none of them can stand properly when the original objective is unclear. Prompt engineering is not becoming obsolete; it is becoming integrated into AI systems engineering.


    🧩 記事の背景

    🧩 01|English Background

    AI agents now perform long, tool-based tasks rather than single chat responses. This shift has moved attention from clever wording to information management, execution control, verification, memory, and recurring automation.

    🧩 02|日本語背景

    AIが一度だけ回答するチャットから、複数のツールを使って長時間働くエージェントへ進化したことで、指示文だけでなく、情報管理、権限、検証、記憶、停止条件を含むシステム設計が必要になりました。


    🧩 ハッシュタグ

    #生成AI #AIエージェント #プロンプトエンジニアリング #コンテキストエンジニアリング #ハーネスエンジニアリング #ループエンジニアリング #自律型AI #業務自動化 #ソフトウェア開発 #AI開発 #AI時代 #デジタル人材 #DX #生産性向上 #評価設計 #検証 #ガバナンス #コスト管理 #未来の働き方 #英語学習

    #GenerativeAI #AIAgents #PromptEngineering #ContextEngineering #HarnessEngineering #LoopEngineering #AgenticAI #Automation #SoftwareEngineering #AIEngineering #LLM #RAG #Evaluation #Verification #FutureOfWork #TechTrends

    #教育 #英会話 #習い事

    #2026 /07/02 #2026 #2026 /07


    🧩↓👍イイネを押してもらえると嬉しいです


    🧩 語彙ダジャレ記憶

    🧩01|phrasing

    読み:フレイジング
    意味:言い回し、表現方法
    ダジャレ:「このフレーズ、いいんぐ?」と確認してphrasingを改善!

    🧩02|context

    読み:コンテクスト
    意味:文脈、背景情報
    ダジャレ:「こんなテキスト?」だけでは分からない。前後の文脈がcontext!

    🧩03|finite

    読み:ファイナイト
    意味:有限の、限りがある
    ダジャレ:「ファイトはしても、夜は有限(ファイナイト)」と覚える!

    🧩04|curate

    読み:キュレイト
    意味:情報を選び、整理する
    ダジャレ:「急にレートを見て厳選!」がcurate!

    🧩05|harness

    読み:ハーネス
    意味:制御する仕組み、実行基盤
    ダジャレ:「はあ、寝ずに暴走?」を防いで制御するharness!

    🧩06|persistent

    読み:パーシステント
    意味:持続する、永続的な
    ダジャレ:「パーッと消えず、ずっとステイ」するpersistent!

    🧩07|verify

    読み:ヴェリファイ
    意味:検証する、正しいか確かめる
    ダジャレ:「Very fineなの?」とverifyして確かめる!

    🧩08|autonomous

    読み:オートノマス
    意味:自律的な、自動で行動する
    ダジャレ:「オートで任す」がautonomous!

    🧩09|trigger

    読み:トリガー
    意味:きっかけ、作動させるもの
    ダジャレ:「鳥がガーッと鳴いたら開始!」その合図がtrigger!

    🧩10|termination

    読み:ターミネーション
    意味:終了、停止
    ダジャレ:「ターミナルで終了!」とterminationを覚える!

    🧩11|constraint

    読み:コンストレイント
    意味:制約、制限条件
    ダジャレ:「このストレートには制約あり!」がconstraint!

    🧩12|bottleneck

    読み:ボトルネック
    意味:進行を妨げる要因
    ダジャレ:ボトルのneckが細いと流れが詰まる。そこがbottleneck!




     
     
    世界のDailyNewsArticleの紹介です。英語を勉強するに際し面白そうなものを選びました。

    あなたへのおすすめ