AIAgentsAI CodingInfrastructureChina AI
AI 日报 | 2026-09-28
北京时间 2026-09-28 · 近 24–48 小时重点更新
今日概览
过去 24–48 小时没有出现单一“改写格局”的基础模型发布,但企业 Agent、推理基础设施和算力商业化继续推进。今天重点看五条:Microsoft 把 Copilot 扩展成 Home/Code/Autopilot 三层工作入口;AWS 推出 SageMaker HyperPod Inference Gateway;Claude Code 继续补齐企业级 MCP、权限、模型治理;阿里 Qwen Intelligence 把通义能力打包给手机厂商;Modal Labs 据 Bloomberg/Reuters 报道正洽谈以约 150 亿美元估值融资。
最重要 5 条
1. Microsoft 发布新 Copilot:Home、Code、Autopilot,把“办公助手”推向 Agent 平台
摘要:Microsoft 9 月 25 日宣布新 Copilot,核心是 Home、Code、Autopilot。Code 让非专业开发者在 Copilot 中构建解决方案,底层使用 GitHub Copilot 技术;Autopilot 是可持续工作的个人 Agent;Copilot Managed Runtime 提供在 Microsoft 365 租户内安全托管代码与应用的运行时。
关键细节:Code 将在 Frontier 计划中推出;Managed Runtime 已进入 preview;Fabric IQ、Dynamics 365、Power Platform 数据会成为 Copilot grounding;Agent 365 成本管理覆盖 Code 与 Runtime。
为什么重要:这不是单个聊天功能更新,而是 Microsoft 把 M365、GitHub Copilot、Fabric、Power Platform 和 Agent 365 成本治理连成企业 Agent 操作系统。
来源:https://blogs.microsoft.com/blog/2026/09/25/introducing-the-new-copilot-with-home-code-and-autopilot/
2. AWS 推出 SageMaker HyperPod Inference Gateway:面向 LLM 推理的 GPU-aware 路由
摘要:AWS 9 月 24 日发布 SageMaker HyperPod Inference Gateway,一个 Kubernetes-native、GPU-aware 的推理路由系统,可作为 EKS managed add-on 部署在 SageMaker HyperPod 上。
关键细节:Endpoint Picker 会实时使用 KV cache utilization、queue depth、LoRA adapter residency、prefix cache hit rate、predicted latency、running requests 等 6 类信号选择 pod;AWS 称 first-token latency 可降低最高 82%,p99 TTFT 降低 97–98%。
为什么重要:LLM 推理瓶颈正在从有没有 GPU转向如何调度 KV cache、LoRA、prefix cache、队列和异构 GPU 池。
来源:https://aws.amazon.com/about-aws/whats-new/2026/09/sagemaker-hyperpod-inference-gateway/
3. Claude Code 继续企业化:AGENTS.md、Auto mode、模型 allow/deny、MCP 与插件稳定性
摘要:Claude Code 9 月下旬更新重点是企业级 coding agent 的治理和运维,包括 AGENTS.md fallback、Auto mode server-side classifier、模型白名单/黑名单、prompt audit、MCP/插件稳定性修复。
关键细节:近期版本加入 /doctor prompt-audit;deniedModels 与 availableModelsMatch: exact 控制模型可用范围;MCP URL-mode elicitation、managed MCP、gateway hint header 强化组织级接入。
为什么重要:AI coding 的下一阶段是可托管、可审计、可批量运行;企业卡点在权限、上下文规范、工具链治理、成本和审计链路。
来源:https://code.claude.com/docs/en/whats-new/
4. 阿里 Qwen Intelligence 面向手机厂商:Agentic smartphones 的全栈方案
摘要:阿里 9 月 25 日披露 Qwen Intelligence,全栈 Agentic AI 方案,面向手机厂商提供从手机优化基础模型到 ready-to-use agents 的能力。
关键细节:HONOR Magic9 Series 与 HONOR Robot Phone 将作为首批设备接入;方案组合通义模型能力、手机场景 Agent、设备能力调用和厂商体验层。
为什么重要:中国 AI 生态落点正在从 Web chatbot/API 转到硬件入口;模型厂商争夺的是默认系统智能层。
来源:https://www.alibabacloud.com/blog/alibaba-launches-qwen-intelligence-to-power-next-generation-agentic-smartphones_603597
5. Modal Labs 据报洽谈以约 150 亿美元估值融资
摘要:Reuters 援引 Bloomberg 报道,AI startup Modal Labs 正洽谈新融资,估值约 150 亿美元。
关键细节:Modal 位于 AI/数据/批处理工作负载的 serverless compute 赛道,连接开发者体验、GPU/CPU 弹性算力和作业编排。
为什么重要:资本市场仍在奖励介于 hyperscaler 与应用公司之间的 AI infra 抽象层;Agent、eval、RL、batch inference 增长会放大这种需求。
来源:https://www.reuters.com/technology/modal-labs-talks-raise-funds-15-billion-valuation-bloomberg-news-reports-2026-09-23
其他值得关注
- NVIDIA Vera Rubin NVL72 在 MLPerf Inference v6.1 中首次亮相,继续把竞争焦点拉到 rack-scale、NVLink、推理吞吐与能效。
- CoreWeave 9 月中旬宣布多机架 NVIDIA Vera Rubin NVL72 集群上线,并强调跨区域写加速与 AI Object Storage。
- Qwen Code 9 月 17 日周报显示其可把子任务 hand off 给 Claude Code 和 Codex,并支持命名 workflow、token cap、Goal turn/time 限制。
- Qwen3.8-Omni-Flash 与 Qwen-Image-2.1 在 9 月下旬继续扩展多模态与图像生成/编辑能力。
- 中国模型侧过去 48 小时未检索到 DeepSeek、Kimi、GLM、豆包、文心、MiniMax 的重大官方新发布;今日保留观察,不用二手传闻填充。
来源链接
- https://blogs.microsoft.com/blog/2026/09/25/introducing-the-new-copilot-with-home-code-and-autopilot/
- https://aws.amazon.com/about-aws/whats-new/2026/09/sagemaker-hyperpod-inference-gateway/
- https://code.claude.com/docs/en/whats-new/
- https://www.alibabacloud.com/blog/alibaba-launches-qwen-intelligence-to-power-next-generation-agentic-smartphones_603597
- https://www.reuters.com/technology/modal-labs-talks-raise-funds-15-billion-valuation-bloomberg-news-reports-2026-09-23
- https://qwenlm.github.io/qwen-code-docs/en/blog/updates/weekly-update-2026-09-17/
- https://www.alibabacloud.com/blog/qwen3-8-omni-flash-omni-senses--agentic-delivery-_603580
- https://www.alibabacloud.com/blog/qwen-image-2-1-compact-efficient-and-unified-image-creation_603586
- https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/