穗稻忠武的專欄
HF 速讀AI 論文

HF 論文速讀 2026-06-25

HF 論文速讀 2026-06-25

今天的共同趨勢:AI 研究正在從模型能力本身,推進到長任務、長上下文、空間感知與可驗證環境。

今天先追 NatureBench:Coding Agent 離科學發現還差一段路,再用另外四篇補齊 Agent、robotics、spatial reasoning、context efficiency 與 SOTA 感知能力的路線圖。

今天先記三件事

  1. 先看是否有 SOTA / 榜單訊號,再看它能不能落地到產品或工程流程。
  2. Agent、長上下文、世界模型、空間推理正在變成同一件事:讓模型更可靠地做長任務。
  3. 今天的 50 頁簡報可以當附錄;真正要先記的是每篇 paper 解決的瓶頸與適用場景。

五篇速讀

NatureBench:Coding Agent 離科學發現還差一段路

LLM/Agent 基…

為什麼重要:We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-re…

帶走什麼:We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-re…

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

Nemotron 3 Nano Omni:3B active 的開放多模態模型

LLM/Agent 基…

為什麼重要:OCRBench-V2 EN; reasoning on; source: paper Table 7; MoE total=30B active=3B; arXiv paper…

帶走什麼:OCRBench-V2 EN; reasoning on; source: paper Table 7; MoE total=30B active=3B; arXiv paper…

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

MinerU2.5-Pro:文件解析開始走向 data-centric scale

LLM/Agent 基…

為什麼重要:ParseBench overall score from the comparison table in ParseBench: A Document Parsing Benc…

帶走什麼:ParseBench overall score from the comparison table in ParseBench: A Document Parsing Benc…

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

BoostTrack++:用 tracklet 找回多目標追蹤漏掉的人

研究趨勢

為什麼重要:MOT20 test set, private detection protocol.

帶走什麼:MOT20 test set, private detection protocol.

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。

MemGUI-Agent:長任務 Mobile GUI Agent 需要主動整理記憶

LLM/Agent 基…

為什麼重要:MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet r…

帶走什麼:MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet r…

下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。