HF 論文速讀 2026-06-25

今天的共同趨勢:AI 研究正在從模型能力本身,推進到長任務、長上下文、空間感知與可驗證環境。
今天先追 NatureBench:Coding Agent 離科學發現還差一段路,再用另外四篇補齊 Agent、robotics、spatial reasoning、context efficiency 與 SOTA 感知能力的路線圖。
今天先記三件事
- 先看是否有 SOTA / 榜單訊號,再看它能不能落地到產品或工程流程。
- Agent、長上下文、世界模型、空間推理正在變成同一件事:讓模型更可靠地做長任務。
- 今天的 50 頁簡報可以當附錄;真正要先記的是每篇 paper 解決的瓶頸與適用場景。
五篇速讀
NatureBench:Coding Agent 離科學發現還差一段路
LLM/Agent 基…
為什麼重要:We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-re…
帶走什麼:We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-re…
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
Nemotron 3 Nano Omni:3B active 的開放多模態模型
LLM/Agent 基…
為什麼重要:OCRBench-V2 EN; reasoning on; source: paper Table 7; MoE total=30B active=3B; arXiv paper…
帶走什麼:OCRBench-V2 EN; reasoning on; source: paper Table 7; MoE total=30B active=3B; arXiv paper…
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
MinerU2.5-Pro:文件解析開始走向 data-centric scale
LLM/Agent 基…
為什麼重要:ParseBench overall score from the comparison table in ParseBench: A Document Parsing Benc…
帶走什麼:ParseBench overall score from the comparison table in ParseBench: A Document Parsing Benc…
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
BoostTrack++:用 tracklet 找回多目標追蹤漏掉的人
研究趨勢
為什麼重要:MOT20 test set, private detection protocol.
帶走什麼:MOT20 test set, private detection protocol.
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。
MemGUI-Agent:長任務 Mobile GUI Agent 需要主動整理記憶
LLM/Agent 基…
為什麼重要:MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet r…
帶走什麼:MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet r…
下一步:先看問題設定與 benchmark,再決定是否需要讀完整簡報。