HF 論文速讀 2026-07-31

今天的共同趨勢:AI 研究正在從模型能力本身,推進到長任務、長上下文、空間感知與可驗證環境。
今天先追 S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation,再用另外四篇補齊 Agent、roboti…
今天先記三件事
- 先看是否有 SOTA / 榜單訊號,再看它能不能落地到產品或工程流程。
- Agent、長上下文、世界模型、空間推理正在變成同一件事:讓模型更可靠地做長任務。
- 今天的 50 頁簡報可以當附錄;真正要先記的是每篇 paper 解決的瓶頸與適用場景。
五篇速讀
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Pre…
多模態生成
為什麼重要:Paper Table 29; CVTG-2K text-rendering evaluation.
帶走什麼:Paper Table 29; CVTG-2K text-rendering evaluation.
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
Loop the Loopies!
LLM/Agent 基…
為什麼重要:Table 3 (OlympiadBench); evaluated with EvalScope.
帶走什麼:Table 3 (OlympiadBench); evaluated with EvalScope.
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
GRASP: GRanularity-Aware Search Policy for Agentic RAG
LLM/Agent 基…
為什麼重要:Table 2; HotpotQA validation evaluation subset, token-level F1.
帶走什麼:Table 2; HotpotQA validation evaluation subset, token-level F1.
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Vide…
研究趨勢
為什麼重要:Table 1 (OmniBench); best result reported for the paper-introduced AVF model family.
帶走什麼:Table 1 (OmniBench); best result reported for the paper-introduced AVF model family.
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
研究趨勢
為什麼重要:Table 1; AIME 2024 pass@1 accuracy averaged over 64 runs, High inference mode (128K trunc…
帶走什麼:Table 1; AIME 2024 pass@1 accuracy averaged over 64 runs, High inference mode (128K trunc…
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。