很多 agent demo 看起來會成功,是因為任務很短。
模型讀一段需求,呼叫一兩個工具,最後回一段像樣的答案。這種情境下,只要 SDK 能發訊息、能 tool calling、能 streaming,通常就足夠展示效果。真正麻煩的是另一種任務:它要查很多資料、分派子任務、讀寫檔案、跑 shell、保留工作脈絡、遇到危險操作要停下來問人,最後還要能被追蹤、部署和重跑。
這時候問題就不再只是「我要不要做 agent」,而是「這個 agent 有沒有一副能撐長任務的骨架」。
Deep Agents 值得看的原因,就在這裡。它不是 LangChain 生態裡又多一個簡單 agent wrapper,而是把長任務 agent 常見需要的 harness 能力預先包在一起:subagents、filesystem、context management、shell access、persistent memory、human-in-the-loop、skills、MCP tools,以及建立在 LangGraph 上的 streaming、persistence、checkpointing。
English TL;DR
- Deep Agents is LangChain’s open-source, batteries-included agent harness for long-horizon, multi-step agents.
- Its value is bundling the operational pieces that serious agents quickly need: subagents, filesystem access, context management, memory, approvals, skills, tools, tracing, and production paths through LangGraph/LangSmith.
- It is useful for engineering teams building research agents, coding agents, internal workflow agents, or agent products that must run beyond a single turn.
- It is not a product strategy shortcut. You still need tool boundaries, sandboxing, evals, user experience, authorization, and failure handling.
- Pragmatic takeaway: evaluate Deep Agents when your agent problem has become harness engineering, not when you are still validating whether the use case deserves an agent.
Deep Agents 是什麼
Deep Agents 是 langchain-ai/deepagents 維護的開源 agent harness。官方 README 把它定位成「batteries-included agent harness」,也就是一個有意見、可擴充、model-agnostic、偏 production-ready 的 agent 骨架。它建立在 LangGraph 之上,和 LangChain 的 create_agent、LangGraph 的自訂 graph 並不是同一層。
比較務實的理解是這樣:
- LangChain
create_agent比較像輕量 agent 起點。 - LangGraph 是可以自訂狀態、流程與節點的 runtime。
- Deep Agents 則是在這兩者上面先包一組長任務 agent 常常會需要的預設能力。
所以它的價值不是「讓模型變聰明」。它真正想省掉的,是每個團隊一旦開始做 serious agent,最後都會重複補的那堆膠水:工作目錄怎麼管理、工具輸出太大怎麼處理、子任務怎麼拆、記憶怎麼留、人類審核怎麼插進 loop、跨 session 怎麼續、上線後怎麼 trace。
截至 2026-09-01 早上查 GitHub API,langchain-ai/deepagents 約有 28.7k stars,license 是 MIT,repo 在 2026-08-31 UTC 仍有 push。最近 release 包含 deepagents==0.7.11、deepagents-code==0.1.65 等 2026-08-28 發布的版本,recent commits 也還在處理 SDK、Talon channel、agent activity logging、rubric grader hooks 這類偏實作與營運的問題。這不是只停在概念頁的 repo,而是正在快速把「agent harness」這層做厚。
為什麼現在值得看
現在很多 agent 專案正在從「單輪工具呼叫」走向「長時間任務」。這個變化很關鍵,因為長任務的失敗模式完全不同。
短任務失敗,多半是答案不準、工具參數錯、格式壞掉。長任務失敗,會變成另一種東西:context 被中間結果塞滿、模型忘記主要目標、工具輸出把重要資訊淹掉、子任務互相干擾、危險操作沒有 approval、記憶跨使用者污染、跑一半中斷後無法恢復,最後連工程師也不知道它到底在哪一步走歪。
Deep Agents 的切角剛好是這些問題。官方文件的 subagents 頁明確把 subagents 用來處理 context bloat:讓子 agent 處理大量 tool calls,只把最後結果回傳給主 agent。human-in-the-loop 文件則把敏感工具操作拉進 interrupt flow,可以 approve、edit、reject 或 respond。production 文件更直接談 thread、user、assistant 這三種 scope,以及 production 裡要面對的 invocation、multi-tenancy、authentication、credentials、async、durability、memory、execution environment、guardrails。
這些都不是 demo 會炫的功能,但它們是 agent 真的開始碰業務流程時,會很快變成主菜的問題。
適合誰
第一種適合的人,是正在做內部 research agent 或分析助理的工程團隊。這類 agent 不只是回答一句話,而是會搜尋資料、讀文件、整理中間結果、可能產出檔案或報告。Deep Agents 的 filesystem、context management、subagents 和 tracing,在這裡會比單純 tool calling 更有意義。
第二種,是要做 coding agent 或 repo 內部助理的團隊。它需要讀檔、改檔、跑測試、呼叫 shell、分派 code review 或 research 子任務,甚至接 skills。Deep Agents 本身也延伸出 Deep Agents Code,官方 README 直接把它描述成類似 Claude Code 或 Cursor 的 terminal coding agent。這代表它不只是抽象 SDK,也正在吃自己那套 agent harness 的產品形狀。
第三種,是已經有 LangChain / LangGraph / LangSmith 投資的團隊。Deep Agents 的一大好處,是不用把既有 stack 全部推倒。你可以把它當成更高層的 harness,底下仍然吃 LangGraph 的 persistence、checkpointing、streaming,上線時也能接 LangSmith tracing、evaluation、monitoring 和 deployment。
第四種,是需要可審核工具操作的企業團隊。只要 agent 會刪檔、寄信、改資料、開 PR、跑部署、查內部 observability,human-in-the-loop 就不是錦上添花,而是上線門檻。Deep Agents 的 interrupt 設計至少把這件事放進框架層,而不是要你每次在工具函式裡手刻一套 approval 狀態機。
不適合誰
如果你只是做簡單 chatbot,Deep Agents 很可能太重。單輪 FAQ、客服草稿、內容改寫、簡單分類,通常用 provider SDK、LangChain create_agent、Pydantic AI、Mirascope 或一小段固定 pipeline 就夠了。把完整 harness 拉進來,可能只是讓專案更早背上不需要的抽象。
如果你的流程其實是 deterministic workflow,也不一定要用它。表單進來,查資料庫,套規則,寄通知,更新 CRM,這種任務很多時候應該用 Temporal、Prefect、n8n、Airflow、Cloud Tasks 或後端 service 排程處理。Agent 適合處理語意、判斷與不確定性,不代表所有流程都該 agent 化。
如果你的團隊還沒有工具權限設計,也不該急著上。Deep Agents README 的 security 段講得很直接:它採取的是 trust the LLM model,也就是 agent 能做什麼,取決於工具和 sandbox 邊界,而不是期待模型自我約束。這句話很重要。框架能幫你插 approval、設 middleware、接 sandbox,但它不能替你決定哪些操作本來就不該給 agent。
兩個具體場景
場景一,研究型 agent 從摘要器變成工作台
假設一個投資研究或產品策略團隊想做內部 research agent。最初版本可能只是丟一個問題,讓模型搜尋網頁、整理摘要。這很快就會撞到天花板:資料來源很多,中間筆記很長,不同子題需要不同搜尋策略,最後報告還要能追溯來源。
Deep Agents 的 subagents 在這種場景很自然。主 agent 負責拆題和最後判斷,research subagent 去查市場資料,technical subagent 去讀 repo 或 docs,競品 subagent 去整理替代方案。主 agent 不需要吞下所有 raw tool output,只需要拿到每個子任務的結論。filesystem 則可以用來存中間筆記、表格、草稿與引用,避免所有東西都塞在 context 裡。
這種設計不保證答案一定對,但至少把長任務拆成比較可治理的工作單位。對 research agent 來說,這比單純把 prompt 寫長更接近可維護系統。
場景二,內部 coding agent 需要從「會改檔」走向「能協作」
Coding agent 最容易被低估的地方,是它不只是要會寫 code。它還要知道 repo 規則、能讀 AGENTS.md 或專案文件、能分派子任務、能跑測試、能在危險命令前停下來、能把過程留下 trace,最後還要讓人 review。
Deep Agents 的組合剛好對上這個問題。filesystem 和 shell access 讓 agent 能在受控環境裡工作;skills 讓團隊把常見流程包成可重用能力;subagents 可以把測試、搜尋、review 拆出去;human-in-the-loop 可以把刪檔、push、部署或外部通知變成需要批准的節點。
這裡的重點不是它會不會比 Claude Code 或 OpenCode 更會改 code,而是它更像一個可嵌入產品與內部平台的 harness。若公司要自己做 coding agent,而不是只買現成工具,Deep Agents 會是一個值得研究的底座。
限制與缺陷
第一個限制,是它仍然需要工程能力。Deep Agents 不是 no-code agent builder。你要理解 LangChain model provider、工具定義、middleware、checkpointer、filesystem backend、memory scope、deployment 與 observability。對已經有平台工程能力的團隊,這是彈性;對只是想快速做 demo 的團隊,這就是負擔。
第二個限制,是安全邊界不能外包給 harness。官方文件可以提供 human-in-the-loop、guardrails、production guidance,但真正的權限仍然要由你設計。哪些工具可讀、哪些可寫、哪些需要審核、sandbox 能不能連網、secret 怎麼處理、multi-tenant memory 是否隔離,這些都不是安裝套件後自動變正確。
第三個限制,是 LangChain 生態黏著度。Deep Agents 的優勢正是它建立在 LangGraph / LangSmith / LangChain 之上,但這也代表如果你的團隊本來不在這個生態,導入成本會比純 provider SDK 或小型框架高。你得到的是一組完整度較高的堆疊,也同時接受了它的抽象、版本節奏與平台路線。
第四個限制,是「長任務」本身還是難。Context management、subagents、memory、approval 都只能降低混亂,不會讓模型突然具備完美計畫能力。長任務 agent 仍然會遇到偏題、反覆嘗試、成本超支、錯誤恢復、工具誤用與結果驗收問題。Deep Agents 補的是骨架,不是產品成功保證。
採用判斷
Deep Agents 最適合的採用時機,是你的 agent 已經不再是單一 prompt 或單一 tool call,而開始需要長任務能力。
如果你現在的痛點是:
- agent context 常常爆掉或被中間資料污染
- 需要子任務與不同專長的 agent
- 需要讀寫檔案、跑 shell、產出 artifacts
- 需要敏感操作 approval
- 需要 persistent memory、checkpointing、tracing、production deployment
- 已經在 LangGraph / LangSmith 生態裡
那 Deep Agents 很值得現在看。
但如果你還在驗證 use case、任務很短、工具很少、流程其實固定,較穩健的做法是先用更小的工具把問題跑清楚。等你真的開始感受到 harness engineering 的重量,再導入 Deep Agents 才比較合理。
結論不是「所有 agent 都該用 Deep Agents」。比較準確的說法是:當你的 agent 開始像一個會長時間工作的系統,而不是一次性回覆介面,Deep Agents 才真正對題。
它的價值在於承認 agent productization 不只靠模型,而是靠一整套執行骨架。這副骨架如果自己慢慢補,通常會補成一堆分散的 middleware、工具 wrapper、記憶 hack、approval if/else 和部署腳本。Deep Agents 的野心,就是把這些東西先整理成可延伸的 harness。
但最後仍然要回到最現實的一句話:harness 只決定 agent 怎麼跑,不決定它該不該跑。真正的採用判斷,還是要看任務本身是否值得交給 agent,以及你是否已經準備好管理它的權限、成本、失敗與責任。
GitHub Star History
參考資料
- GitHub Repo: https://github.com/langchain-ai/deepagents
- README: https://github.com/langchain-ai/deepagents/blob/main/README.md
- Deep Agents overview: https://docs.langchain.com/oss/python/deepagents/overview
- Subagents docs: https://docs.langchain.com/oss/python/deepagents/subagents
- Human-in-the-loop docs: https://docs.langchain.com/oss/python/deepagents/human-in-the-loop
- Going to production docs: https://docs.langchain.com/oss/python/deepagents/going-to-production
- GitHub API Metadata: https://api.github.com/repos/langchain-ai/deepagents
- Recent commits: https://api.github.com/repos/langchain-ai/deepagents/commits?per_page=5
- Releases: https://github.com/langchain-ai/deepagents/releases
- Star History: https://www.star-history.com/#langchain-ai/deepagents&Date