穗稻忠武的專欄
HF 速讀AI 論文

HF 論文速讀 2026-07-31

HF 論文速讀 2026-07-31

今天的共同趨勢:AI 研究正在從模型能力本身,推進到長任務、長上下文、空間感知與可驗證環境。

今天先追 S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation,再用另外四篇補齊 Agent、roboti…

今天先記三件事

  1. 先看是否有 SOTA / 榜單訊號,再看它能不能落地到產品或工程流程。
  2. Agent、長上下文、世界模型、空間推理正在變成同一件事:讓模型更可靠地做長任務。
  3. 今天的 50 頁簡報可以當附錄;真正要先記的是每篇 paper 解決的瓶頸與適用場景。

五篇速讀

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Pre…

多模態生成

為什麼重要:Paper Table 29; CVTG-2K text-rendering evaluation.

帶走什麼:Paper Table 29; CVTG-2K text-rendering evaluation.

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

Loop the Loopies!

LLM/Agent 基…

為什麼重要:Table 3 (OlympiadBench); evaluated with EvalScope.

帶走什麼:Table 3 (OlympiadBench); evaluated with EvalScope.

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

GRASP: GRanularity-Aware Search Policy for Agentic RAG

LLM/Agent 基…

為什麼重要:Table 2; HotpotQA validation evaluation subset, token-level F1.

帶走什麼:Table 2; HotpotQA validation evaluation subset, token-level F1.

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Vide…

研究趨勢

為什麼重要:Table 1 (OmniBench); best result reported for the paper-introduced AVF model family.

帶走什麼:Table 1 (OmniBench); best result reported for the paper-introduced AVF model family.

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

研究趨勢

為什麼重要:Table 1; AIME 2024 pass@1 accuracy averaged over 64 runs, High inference mode (128K trunc…

帶走什麼:Table 1; AIME 2024 pass@1 accuracy averaged over 64 runs, High inference mode (128K trunc…

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。