AI 每日快讯

AI 每日快讯

AI 产品、模型、开源工具和官方动态的时间流。保留历史记录,按分类、日期和标签继续筛选。

3779历史快讯
207开源工具
80当前结果
10 月 02 日 今日快讯
MarkTechPost 官方资讯

MarkTechPost:A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired …

原文摘要:In this comprehensive coding guide, we explore Google Research's Kauldron—a JAX training library optimized for research velocity and modularity. Learn how konfig turns experiments 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Proba…

原文摘要:Cloudflare has released Clef (27B) and Clef-flash (9B), open-weight decision models that return typed probabilities instead of text. They are Jev-API compatible, accept images, and 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 30 日 2026-09-30 快讯
MarkTechPost 官方资讯

MarkTechPost:OpenAI Releases GPT-6.1 Sol: Near-Astra Coding and Computer Use at One-Fifth of Astra’s Toke…

原文摘要:OpenAI released GPT-6.1 Sol on September 29, 2026, an upgrade to GPT-6 Sol. It reaches near-Astra results on agentic coding, computer use and professional work at one-fifth of Astr 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and …

原文摘要:Gemini 4 Argon tops GPT-6 Astra and Claude Opus 5.5 on most 评测, but access remains gated today. The post Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Co 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 m…

原文摘要:Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Ph 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos …

原文摘要:Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 29 日 2026-09-29 快讯
MarkTechPost 官方资讯

MarkTechPost:H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Ac…

原文摘要:H Company has released Holo4, a family of generalist computer-use models for AI agents. One set of weights clicks and types on screens. It also writes code and calls MCP or API too 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

原文摘要:Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It scores 70.6% on Terminal-Bench 4.0 and lands within 2 points of Opus 5.5 on GDPval-AA. It al 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 27 日 2026-09-27 快讯
MarkTechPost 官方资讯

MarkTechPost:A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the 评测 Contract a…

原文摘要:A comprehensive coding tutorial on Google Research's Massive Sound Embedding 评测 (MSEB), demonstrating how to implement custom sound encoders, drive classification, clusterin 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 26 日 2026-09-26 快讯
MarkTechPost 官方资讯

MarkTechPost:Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global …

原文摘要:Sarvam AI's Saaras V4 is a speech-to-text model covering all 22 Indian languages plus global English. It pairs an audio encoder with a 3B hybrid state-space decoder. It adds keyter 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Build…

原文摘要:Exa has released Agent Ultra, the highest effort mode of its Exa Agent API. It coordinates subagents across thousands of sources for list building and entity enrichment. Exa report 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 25 日 2026-09-25 快讯
MarkTechPost 官方资讯

MarkTechPost:Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With…

原文摘要:Liquid AI has released LFM2.5-VL-3B-DSpark, a 279.5M-parameter draft model that brings speculative decoding to its LFM2.5-VL-3B vision-language model. It delivers up to 3.13x faste 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 24 日 2026-09-24 快讯
MarkTechPost 官方资讯

MarkTechPost:Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× …

原文摘要:Contrastive-LM has released CLM-8B, an open System One model that scores candidate actions against a state instead of generating text. It adds 2 small projection heads to a frozen 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative F…

原文摘要:This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK, u 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 23 日 2026-09-23 快讯
MarkTechPost 官方资讯

MarkTechPost:Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

原文摘要:Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, 2 new text-to-speech models available now through the Gemini API and Google AI Studio. Flash TTS designs new voices fro 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 22 日 2026-09-22 快讯
MarkTechPost 官方资讯

MarkTechPost:Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Th…

原文摘要:Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the level of Claude Fable 5.1 on most work. It also costs 40% l 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 4…

原文摘要:NVIDIA researchers have released SoL-Pi, 4 harness mechanisms for the open-source Pi coding agent, discovered by an AI running auto-research loops across 535 environments. On EdgeB 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

原文摘要:SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks, and knowledge work. Grok 4.7 is built on a larger base model and a longer reinforcement learning r 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 21 日 2026-09-21 快讯
MarkTechPost 官方资讯

MarkTechPost:AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lowe…

原文摘要:Many 开发者 find that an agent idea works inside Claude Code or Codex, then struggles once they rebuild it with their own loop. The Strands Agents team at AWS is targeting that 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 20 日 2026-09-20 快讯
MarkTechPost 官方资讯

MarkTechPost:Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts…

原文摘要:Qwen has released Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model built on a new Interleave architecture. It cuts average lagging (LAAL) from 2.8 seconds to 2. 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 19 日 2026-09-19 快讯
MarkTechPost 官方资讯

MarkTechPost:TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instea…

原文摘要:TypeSafe AI released Jev, a System One model that answers typed questions with probabilities instead of generating text. Input costs $0.042 per 1M tokens, and output tokens are fre 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model

原文摘要:Linkup Research has released SPARSEUP, an open-source sparse embedding model built on a 149M-parameter ModernBERT backbone. It scores 56.4 nDCG@10 on BEIR-13, which Linkup calls th 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:SpaceXAI Releases Grok Voice Transcribe 2.0: A Speech-to-Text API Claiming 2x Accuracy Over …

原文摘要:SpaceXAI has released Grok Voice Transcribe 2.0, its newest speech-to-text model for batch and streaming audio. The company says it is twice as accurate as version 1.0 at the same 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 18 日 2026-09-18 快讯
MarkTechPost 官方资讯

MarkTechPost:Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding …

原文摘要:Jina AI has released jina-ocr-v1, a visual document parser that converts PDFs, scans, tables, charts and invoices into Markdown. The model has 3.4B total parameters, with about 570 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Or…

原文摘要:Building an AI prototype is easy, but operating autonomous agents at scale requires production-grade tooling. Salesforce Agentforce bridges the gap from "vibe coding" to enterprise 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 17 日 2026-09-17 快讯
MarkTechPost 官方资讯

MarkTechPost:Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

原文摘要:Microsoft's AKS engineering team open-sourced TauGrid on August 28, 2026, packaging the tau CLI, Kueue queueing, KubeRay orchestration, GPU node health monitoring and observability 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Inciden…

原文摘要:OpenAI can disclose misalignment before fixes exist. Its 6 initial reports include fabricated data and leaked API keys. The post OpenAI Releases a Model Misalignment Disclosure Fra 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for …

原文摘要:Google Research has introduced Retrieve-for-Train (R4T), a framework for search that returns coherent, diverse result sets. It trains a fan-out language model with RL once, using g 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up …

原文摘要:Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems at once: value quantization erro 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 16 日 2026-09-16 快讯
MarkTechPost 官方资讯

MarkTechPost:Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models

原文摘要:Nums AI has released Causilo, a pretrained tabular foundation model for classification and regression with a scikit-learn interface. It posts the top TabArena Elo among single mode 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 15 日 2026-09-15 快讯
MarkTechPost 官方资讯

MarkTechPost:Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend

原文摘要:Learn how to leverage NVIDIA’s cuDNN Frontend Graph API to build custom kernel fusions, autotuning engine configurations, FP8-style epilogues, scaled dot-product attention, dynamic 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Ag…

原文摘要:Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. The models execute tools and API calls in the background 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

原文摘要:Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an 开源 harness for standing up public 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 14 日 2026-09-14 快讯
MarkTechPost 官方资讯

MarkTechPost:Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleop…

原文摘要:Reward AI has released OM-1 (Omnibody Model 1), a general-purpose manipulation policy trained entirely on human demonstrations captured with a 7-DoF wearable glove, with no teleope 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Tr…

原文摘要:Sakana AI researchers Jeffrey Seely and Julian Gould introduce Augmented Lagrangian Predictive Coding (PC-ALM), a local-learning alternative to backpropagation. By attaching a Lagr 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot …

原文摘要:NVIDIA has open-sourced OSMO, the Kubernetes-native 工作流 orchestrator it uses internally for Project GR00T, Isaac Lab, and Isaac Sim. OSMO lets robotics teams define training, 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 13 日 2026-09-13 快讯
MarkTechPost 官方资讯

MarkTechPost:A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder Stat…

原文摘要:Yifan Zhang's Recurrent Looped Transformer (RLT) technical report proposes a causal encoder paired with a recurrent decoder that carries its final hidden state and layerwise slidin 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Los…

原文摘要:A shallow agent is an LLM calling tools in a loop, and on long tasks it fails in 2 ways: context overflow and goal loss. This article opens the harness layer that fixes both, with 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Implementation of Machine Learning 工作流 with NVIDIA cuML, RAPIDS, GPU Benchmarking, Exp…

原文摘要:This practical tutorial demonstrates how to build and accelerate machine learning 工作流 using NVIDIA cuML and RAPIDS. It covers GPU environment setup, zero-code scikit-learn ac 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 12 日 2026-09-12 快讯
MarkTechPost 官方资讯

MarkTechPost:Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on Fron…

原文摘要:Cognition, the company behind the Devin coding agent, has released SWE-2, its most capable coding model to date. SWE-2 is post-trained with reinforcement learning from Kimi K3, Moo 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its…

原文摘要:The Fly Language Model (FLM) drives all 166,700 retained neurons and 25.6 million edges of the MaleCNS fruit fly connectome with token embeddings, then adds a small learned correct 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 11 日 2026-09-11 快讯
MarkTechPost 官方资讯

MarkTechPost:Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Ch…

原文摘要:ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI introduce HarnessDev, a 评测 that scores the runnable harness a model builds rather than the answer it returns. Sta 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI G…

原文摘要:Anthropic has published a new plugin evals 工作流 for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compar 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use …

原文摘要:Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 10 日 2026-09-10 快讯
MarkTechPost 官方资讯

MarkTechPost:Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Retu…

原文摘要:Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased di 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

原文摘要:OpenAI has released the Agents API in public beta. It gives 开发者 the same harness and infrastructure that run Codex. OpenAI hosts and maintains the harness. 开发者 run th 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

原文摘要:LandingAI has shipped Agentic Document Extraction Gen2, a rebuild of its document stack on the DPT-3 model family. Chunks are retired in favor of a document, page and block tree. D 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 09 日 2026-09-09 快讯
MarkTechPost 官方资讯

MarkTechPost:Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce…

原文摘要:Google has open-sourced Mantis, a stack-agnostic toolkit of security review skills for AI coding agents. It runs the full vulnerability lifecycle: sweep the code, filter false posi 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 06 日 2026-09-06 快讯
MarkTechPost 官方资讯

MarkTechPost:H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That D…

原文摘要:We look at NeoMME, a family of 260M and 800M bidirectional encoders from H Company. Unlike ColPali-style retrievers, it processes multilingual text tokens and raw 32×32 image patch 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 05 日 2026-09-05 快讯
MarkTechPost 官方资讯

MarkTechPost:GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workf…

原文摘要:We look at Project HydraFusion, GitHub's research preview that treats 工作流 selection as an optimization problem rather than a model picker. We break down the three execution pa 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 03 日 2026-09-03 快讯
MarkTechPost 官方资讯

MarkTechPost:OpenAI Releases GPT-6 Astra: A 1.05M-Context Computer-Use Model Gated Behind a ‘Critical’ Cy…

原文摘要:OpenAI released GPT-6 Astra on September 3, 2026, positioning it as a computer-use flagship rather than a chat model. It reports 72.6% on OSWorld V2-Offline, replaces Codex compact 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Anthropic Released Claude Commerce Agents: An Apache-2.0 Blueprint for Shopping and Merchant…

原文摘要:Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and…

原文摘要:Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact model running on the user's ma 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple S…

原文摘要:Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 02 日 2026-09-02 快讯
MarkTechPost 官方资讯

MarkTechPost:Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across Open…

原文摘要:NVIDIA has released Switchyard, an Apache-2.0 Rust proxy and library for LLM traffic. It decodes requests into provider-neutral types, routes them with passthrough, random, LLM-cla 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, G…

原文摘要:Agentic assistants have a structural problem: the context that makes them useful — deal documents, privileged files, client records — is exactly the context users cannot send to a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 01 日 2026-09-01 快讯
MarkTechPost 官方资讯

MarkTechPost:Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science a…

原文摘要:Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, the same model behind two different safeguard layers. Fable 5.1 is generally available on the Claude API, AWS, Google 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 31 日 2026-08-31 快讯
MarkTechPost 官方资讯

MarkTechPost:Keenable AI Open-Sources NEEDLE: A Live Search 评测 That Rebuilds Its Query Set Every H…

原文摘要:How do you 评测 a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent ca 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 30 日 2026-08-30 快讯
MarkTechPost 官方资讯

MarkTechPost:Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-Firs…

原文摘要:Voice agents fail on latency long before they fail on intelligence. Time to first token is the metric most teams use to choose an inference API, and it is the right starting point 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 29 日 2026-08-29 快讯
MarkTechPost 官方资讯

MarkTechPost:Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Contro…

原文摘要:We look at Gemini Omni 1.1 Flash, Google's production update to its native multimodal video generation and editing model. We break down what changed: scene extension now reads up t 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 27 日 2026-08-27 快讯
MarkTechPost 官方资讯

MarkTechPost:Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B,…

原文摘要:Every agent that writes code needs somewhere to run it, and no two vendors quote the same units. This comparison measures burst cold start across E2B, Daytona, Modal, Cloudflare, a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Conti…

原文摘要:Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead o 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 26 日 2026-08-26 快讯
MarkTechPost 官方资讯

MarkTechPost:IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models

原文摘要:IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B sizes, all under Apache 2.0. Every model exposes a thinking / low-effort / non-thinking 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 24 日 2026-08-24 快讯
MarkTechPost 官方资讯

MarkTechPost:Fastino Releases GLiNER2.5: A Boundary-Prediction Architecture That Removes Span Enumeration…

原文摘要:Fastino released GLiNER2.5, replacing span enumeration with boundary prediction so entity width no longer costs compute. Three Apache 2.0 checkpoints ship at 74M, 194M, and 287M pa 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google Research Introduces ME-POIs: A Mobility-Informed Framework that Adds “How a Place Is …

原文摘要:framework that folds aggregate human movement into text-based place embeddings. Language models describe what a place is; they miss how it is used. ME-POIs encodes each visit as a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 22 日 2026-08-22 快讯
MarkTechPost 官方资讯

MarkTechPost:Decoding AI’s Open-Source Course Maps Three Ways to Run an Agent Loop and the Provider Econo…

原文摘要:Most teams treat ‘which model’ as the important decision. The harness engineering literature keeps pointing somewhere else. In LangChain’s Terminal-Bench experiment, changing only 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 21 日 2026-08-21 快讯
MarkTechPost 官方资讯

MarkTechPost:Building Agentic Document Intelligence Pipelines: Creating Scientific Figures with AutoFigur…

原文摘要:This tutorial explores AutoFigure, a practical toolkit for generating professional scientific figures directly from text descriptions and research papers. We walk through setting u 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 20 日 2026-08-20 快讯
MarkTechPost 官方资讯

MarkTechPost:Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Witho…

原文摘要:Three ~300M drafters bring speculative decoding to LFM2.5, delivering up to 3.18x faster decoding with identical greedy output. The post Liquid AI Releases LFM2.5-DSpark Draft Mode 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 18 日 2026-08-18 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native …

原文摘要:NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an Apache-2.0 project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inference 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Spee…

原文摘要:Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboa 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 17 日 2026-08-17 快讯
MarkTechPost 官方资讯

MarkTechPost:DeepSeek AI Releases DeepSeek Harness in 开发者 Preview: An MIT-Licensed Agent Harness Wh…

原文摘要:DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing. 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 14 日 2026-08-14 快讯
MarkTechPost 官方资讯

MarkTechPost:Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Hori…

原文摘要:Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 13 日 2026-08-13 快讯
MarkTechPost 官方资讯

MarkTechPost:Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Obje…

原文摘要:Liquid AI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model built for on-device deployment. It averages 80.7 on ScreenSpot-v2 and lifts RefCOCO grounding from 57.1 to 8 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Cod…

原文摘要:SpaceXAI released Grok 4.6 on August 12, 2026 — a post-training upgrade over Grok 4.5, not a larger base model. It ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligenc 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 11 日 2026-08-11 快讯
MarkTechPost 官方资讯

MarkTechPost:The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open…

原文摘要:LTX-2.5 brings frontier video generation to local NVIDIA hardware: 6.8-second clips, native multishot, day-one ComfyUI, open weights. The post The Video Production Stack Now Fits o 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 10 日 2026-08-10 快讯
MarkTechPost 官方资讯

MarkTechPost:Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GP…

原文摘要:Meta's Muse Glimmer is a 30B open-weights agentic model under Apache 2.0. It fits 24 GB VRAM and decodes 3.1x faster with DFlash speculation. The post Meta AI Releases Muse Glimmer 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。