推理模型與自主代理 Reasoning Models & Agents
把代理系統拆成可觀察的規劃、工具使用、記憶與驗證環節,再判斷改進究竟發生在哪裡。 Decompose agents into observable planning, tool-use, memory, and verification stages, then locate where an improvement actually happens.
單一資料快照 · 2026-07-20Single data snapshot · 2026-07-20
圖表把官方發表資料對齊本站的十類主題。它適合找研究密集區,不能用來判斷品質、影響力或時間變化。The chart aligns official program data with ten site-defined topics. Use it to locate dense areas of work—not to infer quality, impact, or change over time.
資料快照 2026-07-21。篇數表示論文資料涵蓋,不代表研究品質或重要性。 Snapshot 2026-07-21. Counts show data coverage, not research quality or importance.
從篇數進到方法From counts to methods
每條教學用一個可檢驗的問題,說明如何讀該類論文的方法、實驗與適用範圍。Each lesson uses one testable question to examine methods, experiments, and scope in that area.
把代理系統拆成可觀察的規劃、工具使用、記憶與驗證環節,再判斷改進究竟發生在哪裡。 Decompose agents into observable planning, tool-use, memory, and verification stages, then locate where an improvement actually happens.
用狀態、時間與向量場的共同語言,讀懂擴散(diffusion)、流匹配(flow matching)與離散生成方法。 Use a shared language of state, time, and vector fields to read diffusion, flow matching, and discrete generation work.
從構念、測量、分布偏移與不確定性檢查評測基準,而不是只排列排行榜。 Audit benchmarks through constructs, measurement, shift, and uncertainty—not leaderboard rank alone.
沿著感測、表徵、融合與動作四層,定位影像、語音、影片與機器人方法的真正貢獻。 Trace sensing, representation, fusion, and action layers to locate the real contribution in vision, audio, video, and robotics work.
把延遲、吞吐量、記憶體、品質與開發複雜度放進同一張成本表。 Put latency, throughput, memory, quality, and engineering complexity into one cost ledger.
用假設、保證、適用範圍與失效案例四格,快速判斷定理的解釋範圍。 Use assumptions, guarantee, regime, and failure case to judge a theorem's explanatory range.
把基礎模型拆成資料、目標、架構、調適與評測五層,判斷能力變化究竟來自哪一層。 Decompose foundation models into data, objective, architecture, adaptation, and evaluation layers to locate where capability changes originate.
從不變性、保留資訊、介入測試與分布偏移四個角度,檢查表徵能否應付未見情境。 Audit whether representations support unseen contexts through invariance, information, intervention, and shift.
從測量方式、資料切分、機制、驗證到實際使用,分清預測工具、科學假說與臨床證據。 Separate predictive tools, scientific hypotheses, and clinical evidence through measurement, splits, mechanisms, validation, and deployment.
先說清楚要估計的量、依賴的假設、可識別性與校準方式,再區分觀測預測、因果效果與決策不確定性。 Separate observational prediction, causal effects, and decision uncertainty through estimands, assumptions, identifiability, and calibration.
讀圖原則Reading notes
篇數會受分類規則、命名方式與會議收錄範圍影響。它指出密度,不表示重要性。Counts depend on taxonomy, naming, and venue scope. They indicate density, not importance.
判讀新方法時,還要比較基線、消融實驗、計算成本與不確定性。Reading a new method also requires baselines, ablations, compute cost, and uncertainty.
Oral 與 Spotlight 是官方發表資訊,可用於篩選;本站不把它們換算成論文品質分數。Oral and Spotlight are official presentation metadata used for filtering; this site does not convert them into quality scores.
本站不以 spotlight/oral 當作重要性排序;這些欄位只作為可篩選的官方 program metadata。Spotlight and oral labels are never used as importance rankings here; they remain filterable official program metadata.