AI 每日快讯

AI 每日快讯

AI 产品、模型、开源工具和官方动态的时间流。保留历史记录,按分类、日期和标签继续筛选。

3779历史快讯
207开源工具
80当前结果
09 月 30 日 2026-09-30 快讯
MarkTechPost 官方资讯

MarkTechPost:Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and …

原文摘要:Gemini 4 Argon tops GPT-6 Astra and Claude Opus 5.5 on most 评测, but access remains gated today. The post Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Co 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 m…

原文摘要:Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Ph 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos …

原文摘要:Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 28 日 2026-09-28 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA Launches Open Agent Safety Platform: OpenShell Sandboxes Agents on Vera CPUs While Se…

原文摘要:NVIDIA has launched the Open Agent Safety Platform, an open reference design that enforces AI agent safety outside the agent itself. OpenShell, an Apache 2.0 runtime, sandboxes age 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 27 日 2026-09-27 快讯
MarkTechPost 官方资讯

MarkTechPost:A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the 评测 Contract a…

原文摘要:A comprehensive coding tutorial on Google Research's Massive Sound Embedding 评测 (MSEB), demonstrating how to implement custom sound encoders, drive classification, clusterin 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 26 日 2026-09-26 快讯
MarkTechPost 官方资讯

MarkTechPost:Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Build…

原文摘要:Exa has released Agent Ultra, the highest effort mode of its Exa Agent API. It coordinates subagents across thousands of sources for list building and entity enrichment. Exa report 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 24 日 2026-09-24 快讯
MarkTechPost 官方资讯

MarkTechPost:BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accur…

原文摘要:BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that spends 37.2% fewer thinking tokens across 12 评测. Macro accuracy moves from 86.65% to 85.7 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 23 日 2026-09-23 快讯
MarkTechPost 官方资讯

MarkTechPost:Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

原文摘要:Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, 2 new text-to-speech models available now through the Gemini API and Google AI Studio. Flash TTS designs new voices fro 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 22 日 2026-09-22 快讯
MarkTechPost 官方资讯

MarkTechPost:Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Th…

原文摘要:Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the level of Claude Fable 5.1 on most work. It also costs 40% l 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 19 日 2026-09-19 快讯
MarkTechPost 官方资讯

MarkTechPost:TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instea…

原文摘要:TypeSafe AI released Jev, a System One model that answers typed questions with probabilities instead of generating text. Input costs $0.042 per 1M tokens, and output tokens are fre 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 17 日 2026-09-17 快讯
MarkTechPost 官方资讯

MarkTechPost:Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for …

原文摘要:Google Research has introduced Retrieve-for-Train (R4T), a framework for search that returns coherent, diverse result sets. It trains a fan-out language model with RL once, using g 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 13 日 2026-09-13 快讯
MarkTechPost 官方资讯

MarkTechPost:Implementation of Machine Learning 工作流 with NVIDIA cuML, RAPIDS, GPU Benchmarking, Exp…

原文摘要:This practical tutorial demonstrates how to build and accelerate machine learning 工作流 using NVIDIA cuML and RAPIDS. It covers GPU environment setup, zero-code scikit-learn ac 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 11 日 2026-09-11 快讯
MarkTechPost 官方资讯

MarkTechPost:Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Ch…

原文摘要:ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI introduce HarnessDev, a 评测 that scores the runnable harness a model builds rather than the answer it returns. Sta 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI G…

原文摘要:Anthropic has published a new plugin evals 工作流 for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compar 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT2…

原文摘要:Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 10 日 2026-09-10 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput an…

原文摘要:NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch. In 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 07 日 2026-09-07 快讯
09 月 06 日 2026-09-06 快讯
MarkTechPost 官方资讯

MarkTechPost:H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That D…

原文摘要:We look at NeoMME, a family of 260M and 800M bidirectional encoders from H Company. Unlike ColPali-style retrievers, it processes multilingual text tokens and raw 32×32 image patch 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluat…

原文摘要:Training and benchmarking a computer-use agent needs four things — agents, environments, traces, and a framework to evaluate and train them — and all four ship in incompatible form 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

原文摘要:Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineeri 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 05 日 2026-09-05 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA Releases Personal AI Router (PAIR): An 开源 Virtual Inference Router that Dist…

原文摘要:We look at NVIDIA Personal AI Router (PAIR), an 开源 virtual inference router that spreads local AI requests across the machines already on a home network. We cover how PAIR 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 03 日 2026-09-03 快讯
MarkTechPost 官方资讯

MarkTechPost:Anthropic Released Claude Commerce Agents: An Apache-2.0 Blueprint for Shopping and Merchant…

原文摘要:Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and…

原文摘要:Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact model running on the user's ma 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 02 日 2026-09-02 快讯
MarkTechPost 官方资讯

MarkTechPost:Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber: One Core Model, Two Ac…

原文摘要:Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2, 2026. Both variants run on the same foundational intelligence, split by safety mitigations rather than m 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

09 月 01 日 2026-09-01 快讯
MarkTechPost 官方资讯

MarkTechPost:Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First…

原文摘要:Speed and accuracy usually pull against each other in text-to-speech. Gradium AI's new default model reports both: an 81.0% human-rated pass rate on 500 hard sentences across five 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 31 日 2026-08-31 快讯
MarkTechPost 官方资讯

MarkTechPost:Keenable AI Open-Sources NEEDLE: A Live Search 评测 That Rebuilds Its Query Set Every H…

原文摘要:How do you 评测 a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent ca 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate T…

原文摘要:Google Research has released TimesFM-3, a 330 million parameter time series foundation model that forecasts multiple related series in a single forward pass. Unlike every TimesFM c 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 30 日 2026-08-30 快讯
MarkTechPost 官方资讯

MarkTechPost:Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-Firs…

原文摘要:Voice agents fail on latency long before they fail on intelligence. Time to first token is the metric most teams use to choose an inference API, and it is the right starting point 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google AI Introduces EnvHarness: A Programmable Layer That Turns Static Agent Environments I…

原文摘要:Google Cloud AI Research, with Washington University in St. Louis and UNC Chapel Hill, has released EnvHarness, an Apache-2.0 layer that turns a static agent 评测 into one tha 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Anthropic Opens a Research Preview of the Model Hardware Standard (MHS): A Shared Specificat…

原文摘要:Anthropic has opened a research preview of the Model Hardware Standard (MHS), a shared driver specification that lets AI agents discover and safely operate physical devices. Instru 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 27 日 2026-08-27 快讯
MarkTechPost 官方资讯

MarkTechPost:Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Conti…

原文摘要:Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead o 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 25 日 2026-08-25 快讯
MarkTechPost 官方资讯

MarkTechPost:Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Mo…

原文摘要:Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. This week, Liquid AI released Pipette. 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 24 日 2026-08-24 快讯
MarkTechPost 官方资讯

MarkTechPost:Fastino Releases GLiNER2.5: A Boundary-Prediction Architecture That Removes Span Enumeration…

原文摘要:Fastino released GLiNER2.5, replacing span enumeration with boundary prediction so entity width no longer costs compute. Three Apache 2.0 checkpoints ship at 74M, 194M, and 287M pa 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 20 日 2026-08-20 快讯
MarkTechPost 官方资讯

MarkTechPost:Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimizati…

原文摘要:This tutorial provides an end-to-end 工作流 for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 13 日 2026-08-13 快讯
MarkTechPost 官方资讯

MarkTechPost:SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Cod…

原文摘要:SpaceXAI released Grok 4.6 on August 12, 2026 — a post-training upgrade over Grok 4.5, not a larger base model. It ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligenc 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 12 日 2026-08-12 快讯
MarkTechPost 官方资讯

MarkTechPost:AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Eva…

原文摘要:Build a custom LLM post-training pipeline using AllenAI’s Open Instruct framework. This comprehensive guide walks through Supervised Fine-Tuning (SFT), Direct Preference Optimizati 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T W…

原文摘要:Object removal models have improved faster than the metrics used to judge them. Diffusion erasers now reconstruct shadows, reflections and occluded structure convincingly, yet PSNR 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 11 日 2026-08-11 快讯
MarkTechPost 官方资讯

MarkTechPost:webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Loc…

原文摘要:webAI has released TwIL-LM, a family of formal-logic models at 1.7B and 3B parameters that translate English into first-order logic and check whether conclusions follow from premis 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 09 日 2026-08-09 快讯
MarkTechPost 官方资讯

MarkTechPost:Top LLM Observability and 评测 Platforms in 2026: Langfuse, LangSmith, Braintrust, Ari…

原文摘要:A verified 2026 comparison of LLM observability platforms covering tracing depth, 评测 capability, production monitoring, and pricing. The post Top LLM Observability and Eval 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 08 日 2026-08-08 快讯
MarkTechPost 官方资讯

MarkTechPost:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Cl…

原文摘要:Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fi 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 07 日 2026-08-07 快讯
MarkTechPost 官方资讯

MarkTechPost:Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Ta…

原文摘要:Microsoft has open sourced code-testing-generator, a polyglot unit-test agent shipping in the MIT-licensed dotnet/skills 代码仓库. It reads a 代码仓库 before writing anything — 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 06 日 2026-08-06 快讯
MarkTechPost 官方资讯

MarkTechPost:Cloudflare Introduces Kitesurf: An Agent-First Web Browser That Runs Entirely in V8 Isolates…

原文摘要:Cloudflare has released Kitesurf, a stateless web browser built specifically for AI agents that runs entirely in V8 isolates on Cloudflare Workers, with no Chromium underneath. The 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 03 日 2026-08-03 快讯
MarkTechPost 官方资讯

MarkTechPost:Evaluating Multimodal Vision Models with Moonshot PerceptionBench Using Robust Data Loading …

原文摘要:In this tutorial, we design an end-to-end 评测 工作流 for PerceptionBench. This multimodal 评测 measures fine-grained visual perception capabilities across tasks such 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies En…

原文摘要:Cogent AI team released Cogent VR-1, a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general coding streng 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the …

原文摘要:Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a 90-query benc 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

08 月 01 日 2026-08-01 快讯
MarkTechPost 官方资讯

MarkTechPost:Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, …

原文摘要:Discover how to optimize transformer workloads using the NVIDIA Transformer Engine. This tutorial guides you through configuring fused GPU kernels, implementing FP8 delayed scaling 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Supabase Releases Evals: an 开源 评测 That Scores Claude Code, Codex and OpenCod…

原文摘要:Supabase has open sourced supabase/evals, an Apache-2.0 评测 and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — buil 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 25 日 2026-07-25 快讯
MarkTechPost 官方资讯

MarkTechPost:Why the OpenAI Agent Broke Into Hugging Face: Reward Hacking, Not Malice, Explained for Engi…

原文摘要:OpenAI disclosed that its own models breached Hugging Face's production infrastructure while taking a public security 评测. The models were not attacking a target — they were 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 22 日 2026-07-22 快讯
MarkTechPost 官方资讯

MarkTechPost:Research-Grade EdgeBench Analysis: AI Agent Benchmarking, Leaderboard Analytics, Scaling Law…

原文摘要:In this tutorial, we explore EdgeBench as a practical 评测 for evaluating advanced AI agents across diverse task categories, runtime environments, and interaction-time budgets 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vul…

原文摘要:Cisco Foundation AI has released Antares, a family of small language models trained to pinpoint where known vulnerabilities live inside a codebase. Antares-1B reaches 0.209 File F1 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weigh…

原文摘要:Poolside has released Laguna S 2.1, a 118B open-weight Mixture-of-Experts coding model with 8B active parameters per token and a 1M-token context. It matches or beats models severa 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 21 日 2026-07-21 快讯
MarkTechPost 官方资讯

MarkTechPost:Validating Distributed LLM Serving 评测 with NVIDIA srt-slurm, SLURM Recipes, Paramete…

原文摘要:In this tutorial, we explore NVIDIA’s srt-slurm framework and learn how we use srtctl to convert declarative YAML configurations into reproducible SLURM 评测 工作流 for dis 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 19 日 2026-07-19 快讯
MarkTechPost 官方资讯

MarkTechPost:Perplexity AI Releases WANDR: An Open 评测 Evaluating Research Agents That Must Search …

原文摘要:Perplexity's WANDR is an open 评测 and 评测 harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each o 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on 评测…

原文摘要:Three open MoE flagships face off on measured intelligence, MIT versus Modified MIT weights, and real serving cost The post Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Sca 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 17 日 2026-07-17 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks …

原文摘要:NVIDIA released Nemotron 3 Embed on July 15 and 16, 2026. The collection has three open checkpoints: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16, and Nemotron-3-Embed-1B-NVF 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 16 日 2026-07-16 快讯
MarkTechPost 官方资讯

MarkTechPost:OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers …

原文摘要:OpenAI trained GPT-Red, an internal-only attacker model, using self-play reinforcement learning against a population of defender LLMs. It beat human red-teamers 84% to 13% on a rep 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Patter SDK Guide to Building a Restaurant Booking Phone Agent with Dynamic Variables, Guardr…

原文摘要:We explore the Patter SDK by building a voice-agent 工作流 for a restaurant booking use case. We define dynamic caller variables, register callable tools for availability, bookin 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 13 日 2026-07-13 快讯
MarkTechPost 官方资讯

MarkTechPost:Skyfall AI Releases MORPHEUS: A Persistent Enterprise Simulation 评测 That Makes Contin…

原文摘要:MORPHEUS from Skyfall AI is a persistent enterprise simulation platform for continual reinforcement learning. It runs worlds that never reset, using parameterisable regime shifts a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Prime Intellect Releases Verifiers v1: Composable Tasksets, Harnesses, and Runtimes for Agen…

原文摘要:Prime Intellect launched verifiers 0.2.0, previewing a rewritten "v1" core under the verifiers.v1 namespace. It splits an environment into a taskset (what), a harness (how), and a 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 10 日 2026-07-10 快讯
MarkTechPost 官方资讯

MarkTechPost:Kyutai Releases MuScriptor: An Open-Weight Decoder-Only Transformer for Multi-Instrument Mus…

原文摘要:MuScriptor is an open-weight, decoder-only Transformer from Kyutai and Mirelo. Trained on 170k real recordings plus 1.45M synthetic MIDIs, it transcribes full multi-instrument mixe 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Tr…

原文摘要:SensorFM, a wearable health foundation model from Google Research, Google DeepMind, and university collaborators. We walk through its ViT-1D masked-autoencoder backbone, pretrained 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 09 日 2026-07-09 快讯
MarkTechPost 官方资讯

MarkTechPost:Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for …

原文摘要:Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 08 日 2026-07-08 快讯
MarkTechPost 官方资讯

MarkTechPost:SpaceXAI Releases Grok 4.5, a Cursor-Trained Model for Coding, Agentic Tasks, and Knowledge …

原文摘要:SpaceXAI released Grok 4.5, a Cursor-trained model for coding, agentic tasks, and knowledge work. It serves at 80 TPS, costs $2/$6 per million tokens, and ranks #1 on Harvey's Lega 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 06 日 2026-07-06 快讯
MarkTechPost 官方资讯

MarkTechPost:Sakana AI Launches Sakana Translate, a Namazu-Powered Japanese–English–Chinese Translation T…

原文摘要:Sakana AI has added Sakana Translate to Sakana Chat. It runs on the Namazu model series. The tool translates bidirectionally across Japanese, English, and Chinese. Three modes ship 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Training Gemma-3 for Structured Mathematical Reasoning with Tunix GRPO, LoRA Adapters, and G…

原文摘要:We build an end-to-end GRPO training 工作流 that teaches Gemma-3 to reason through GSM8K math problems. We prepare the environment, authenticate with Hugging Face, load Gemma-3, 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 05 日 2026-07-05 快讯
MarkTechPost 官方资讯

MarkTechPost:Meituan Releases LongCat-2.0: A 1.6T-Parameter Open MoE Model with Native 1M Context and Lon…

原文摘要:Meituan has released LongCat-2.0, a 1.6 trillion-parameter Mixture-of-Experts model that activates about 48 billion parameters per token. It pairs a native 1-million-token context, 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:LlamaIndex ‘legal-kb’: Agentic Retrieval over Index v2 with retrieve, find, read, and grep T…

原文摘要:LlamaIndex’s legal-kb is a public reference app that gives agents filesystem-style access to a document knowledge base on Index v2. It exposes retrieve (hybrid semantic search), fi 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 04 日 2026-07-04 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA HORIZON: A Hands-Free Agent that Evolves Git Worktrees and Hits 100% RTL 评测 Co…

原文摘要:A hands-free NVIDIA agent framework hosts each RTL problem as a versioned 代码仓库, reaching 100% completion across 评测. The post NVIDIA HORIZON: A Hands-Free Agent that E 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 03 日 2026-07-03 快讯
MarkTechPost 官方资讯

MarkTechPost:Mistral AI Releases Leanstral 1.5: An Apache-2.0 Lean 4 Code Agent Model Solving 587 of 672 …

原文摘要:Mistral AI released Leanstral 1.5, a free Apache-2.0 code agent model for Lean 4. It saturates miniF2F and solves 587 of 672 PutnamBench problems. The 119B mixture-of-experts activ 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

07 月 01 日 2026-07-01 快讯
MarkTechPost 官方资讯

MarkTechPost:Using Lift to Turn Research PDFs into Structured JSON with Controlled, Schema-Guided Field-L…

原文摘要:In this tutorial, we build a full PDF-to-structured-data 工作流 around Lift, built for controlled 评测 rather than a one-off demo. We prepare a Colab GPU environment, load 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:Anthropic Redeploys Claude Fable 5 on July 1 After US Export Controls Lift, Adds New Cyberse…

原文摘要:Anthropic is redeploying Claude Fable 5 on July 1 after US export controls were lifted. A new safety classifier blocks the technique in the Amazon report over 99% of the time, rout 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

06 月 29 日 2026-06-29 快讯
MarkTechPost 官方资讯

MarkTechPost:NVIDIA BioNeMo Agent Toolkit Turns Biomolecular Models Into Callable Skills for AI Agents in…

原文摘要:NVIDIA's open-source BioNeMo Agent Toolkit turns biomolecular models like OpenFold3, DiffDock, and GenMol into documented, callable skills for AI agents. Each skill describes a mod 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

06 月 26 日 2026-06-26 快讯
MarkTechPost 官方资讯

MarkTechPost:Cursor Study Finds Reward Hacking Inflates Coding-Agent 评测 Scores on SWE-bench Pro

原文摘要:A Cursor study shows coding agents retrieve known fixes instead of deriving them, inflating SWE-bench Pro scores through runtime contamination. The post Cursor Study Finds Reward H 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

06 月 23 日 2026-06-23 快讯
MarkTechPost 官方资讯

MarkTechPost:Datalab Releases lift: A 9B Open-Weights Vision Model That Extracts Structured JSON From PDF…

原文摘要:Datalab released lift, a 9B open-weights vision model that turns PDFs and images into schema-matching JSON. It uses schema-constrained decoding for valid structure and trained abst 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

MarkTechPost 官方资讯

MarkTechPost:How to Use NVIDIA Canary-1B-v2 for ASR, Translation, and Automatic SRT Subtitle Export in Py…

原文摘要:In this tutorial, we build a multilingual ASR and speech translation pipeline with NVIDIA Canary-1B-v2. We load the model on a GPU-enabled runtime, prepare audio into 16 kHz mono, 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。

06 月 22 日 2026-06-22 快讯
MarkTechPost 官方资讯

MarkTechPost:Sakana AI Launches Sakana Fugu: An Orchestration Model That Routes Tasks Across a Swappable …

原文摘要:Fugu and Fugu Ultra route tasks across a swappable model pool, leading most coding, reasoning, and agentic 评测. The post Sakana AI Launches Sakana Fugu: An Orchestration Mod 来源:MarkTechPost。建议继续查看原文,重点核对它影响的工具入口、成本、风险和真实使用场景。