推理模型與自主代理 Reasoning Models & Agents
把代理系統拆成可觀察的規劃、工具使用、記憶與驗證環節,再判斷改進究竟發生在哪裡。 Decompose agents into observable planning, tool-use, memory, and verification stages, then locate where an improvement actually happens.
ICML 20266,628 篇論文6,628 papers資料快照 2026-07-20Snapshot 2026-07-20
先看論文集中在哪些問題,再用摘要、官方主題與來源註記比較個別研究。十條教學路徑補上讀方法、實驗與限制時需要的判斷工具。 See where papers cluster, then compare individual work through abstracts, official topics, and source notes. Ten lessons cover the judgment needed to read methods, experiments, and limitations.
適合想快速掌握會議全貌,或正在找相關工作、基線與研究題目的讀者。具備基本機器學習概念即可開始。For readers surveying the conference or looking for related work, baselines, and research questions. Basic machine-learning knowledge is enough to begin.
閱讀安排Reading plan
選擇時間、目的和主題,系統會列出一組可在時間內完成的閱讀步驟。 Choose a time, goal, and topic to get a set of reading steps that fits the session.
主題分布Topic map
這裡依本站的十個主題統計論文量與摘要覆蓋;方法、證據與限制仍要回到個別論文判讀。 The map reports paper volume and abstract coverage across ten site-defined topics; methods, evidence, and limitations still require paper-level reading.
資料快照 2026-07-21。篇數表示論文資料涵蓋,不代表研究品質或重要性。 Snapshot 2026-07-21. Counts show data coverage, not research quality or importance.
十條研究方法教學Ten research-method lessons
每條約 25–34 分鐘。你會先做預測、比較三種條件、查看解析,再以回想題與應用題收尾。Each takes about 25–34 minutes: make a prediction, compare three conditions, read the explanation, then finish with recall and transfer questions.
把代理系統拆成可觀察的規劃、工具使用、記憶與驗證環節,再判斷改進究竟發生在哪裡。 Decompose agents into observable planning, tool-use, memory, and verification stages, then locate where an improvement actually happens.
用狀態、時間與向量場的共同語言,讀懂擴散(diffusion)、流匹配(flow matching)與離散生成方法。 Use a shared language of state, time, and vector fields to read diffusion, flow matching, and discrete generation work.
從構念、測量、分布偏移與不確定性檢查評測基準,而不是只排列排行榜。 Audit benchmarks through constructs, measurement, shift, and uncertainty—not leaderboard rank alone.
沿著感測、表徵、融合與動作四層,定位影像、語音、影片與機器人方法的真正貢獻。 Trace sensing, representation, fusion, and action layers to locate the real contribution in vision, audio, video, and robotics work.
把延遲、吞吐量、記憶體、品質與開發複雜度放進同一張成本表。 Put latency, throughput, memory, quality, and engineering complexity into one cost ledger.
用假設、保證、適用範圍與失效案例四格,快速判斷定理的解釋範圍。 Use assumptions, guarantee, regime, and failure case to judge a theorem's explanatory range.
把基礎模型拆成資料、目標、架構、調適與評測五層,判斷能力變化究竟來自哪一層。 Decompose foundation models into data, objective, architecture, adaptation, and evaluation layers to locate where capability changes originate.
從不變性、保留資訊、介入測試與分布偏移四個角度,檢查表徵能否應付未見情境。 Audit whether representations support unseen contexts through invariance, information, intervention, and shift.
從測量方式、資料切分、機制、驗證到實際使用,分清預測工具、科學假說與臨床證據。 Separate predictive tools, scientific hypotheses, and clinical evidence through measurement, splits, mechanisms, validation, and deployment.
先說清楚要估計的量、依賴的假設、可識別性與校準方式,再區分觀測預測、因果效果與決策不確定性。 Separate observational prediction, causal effects, and decision uncertainty through estimands, assumptions, identifiability, and calibration.
整理方法How the guide is compiled
開始閱讀Start reading