每日 AI 简报

2026-08-02(内容获取于 08/02 04:43)

AI突破未解数学难题,数学界态度复杂

The Decoder · 08/02 00:01

OpenAI对「单位距离猜想」的反驳引发AI辅助数学进步的浪潮。菲尔兹奖得主蒂莫西·高尔斯称GPT 5.6 Pro解决了其长期研究的问题,但同时警告了AI对数学研究的潜在负面影响。

推荐理由:这篇报道揭示了AI在科学研究前沿的强大能力和潜在伦理挑战,对理解AI未来发展方向至关重要。

GitHub Copilot SDK发布:将AI编程助手集成到应用中

GitHub Trending

GitHub Copilot SDK 是一个多平台开发工具包,旨在帮助开发者将 GitHub Copilot Agent 的AI编程能力集成到自己的应用程序和服务中。它赋能开发者在更广泛的场景中利用AI提升生产力,实现更智能、高效的编程体验。

推荐理由:该SDK为开发者提供了将先进AI编程能力嵌入各类应用的官方途径,对提升开发效率和扩展AI应用边界具重要意义。

法院驳回xAI请求,明州「裸体化」应用禁令生效

TechCrunch · 08/02 04:26

明尼苏达州针对允许用户「裸体化」图像的应用程序禁令将继续生效,此前法院驳回了xAI公司要求阻止该禁令的诉讼请求,凸显了AI伦理与法律监管的冲突。

推荐理由:此案反映了AI应用在伦理和法律层面面临的挑战,对理解未来AI监管趋势具有参考价值。

Reddit CEO质疑谷歌AI概览价值,合作前景不明

Ars Technica · 08/01 20:30

Reddit股价下跌之际,其CEO质疑谷歌AI概览的价值,并表示公司仍在寻求双方的「双赢」合作。这可能预示Reddit与谷歌之间的数据授权合作面临变数。

推荐理由:此新闻揭示了大型AI公司与内容提供商之间复杂的商业利益关系,对于理解AI内容生态的演变有重要参考。

AI邮件跟进代理NudgeForMe,助您把握邮件机会

Product Hunt · 08/02 04:37

NudgeForMe 是一款AI邮件跟进代理,旨在帮助用户处理错过的邮件机会。它能及时提醒重要信息,确保邮件得到关注和回复,从而有效提升沟通效率和商业机会。

推荐理由:对于经常错过重要邮件的用户,NudgeForMe提供了一个简单有效的AI解决方案,值得尝试以提升工作效率。

ChatGPT新增语音控制电脑功能,AI智能体再进化

X 推文 (AttentionVC) · 07/31 19:58

OpenAI推出ChatGPT新功能,用户可与模型进行自然语音对话,同时让它操作电脑。这标志着AI智能体技术又向前迈进了一步,增强了人机交互的深度和广度,预示了更高级别自动化应用的可能。

推荐理由:这一功能显著提升了AI智能体的实用性和交互性,为用户带来了更直观的AI操作体验,预示着未来AI助手的强大潜力。

自动驾驶普及后:深远影响与未来情景分析

X 创作者 (AttentionVC) · 08/01 00:25

该深度探讨分析了自动驾驶汽车普及后的潜在影响,涵盖技术挑战、社会伦理、交通模式改变以及法律法规等多个维度,全面审视了自动驾驶的深远影响。

推荐理由:该分析为理解自动驾驶技术如何重塑社会提供了宝贵视角,对于规划未来战略具有参考意义。

Judge denies xAI’s request to block Minnesota ban on ‘nudify’ apps

Despite a lawsuit from xAI, a Minnesota ban on apps that allow users to “nudify” images can move forward.

中文介绍 明尼苏达州针对允许用户「裸体化」图像的应用程序禁令将继续生效,此前法院驳回了xAI公司要求阻止该禁令的诉讼请求。

Angela Nissel faces down grief with a laugh

Angela Nissel smiles through the pain. | Image: Angela Nissel Angela Nissel's latest book, Good Grief, Pass the Bread, Mom Is Dead, is my kind of memoir. Sure, it's a deeply emotional tale about caring for a terminally ill parent. But it's delivered with the sort of gallows humor that I often turn t

中文介绍 安吉拉·尼塞尔(Angela Nissel)通过其最新回忆录《Good Grief, Pass the Bread, Mom Is Dead》以幽默的方式面对失去身患绝症父母的悲痛。

YouTuber Hank Green says his AI usage is ‘not healthy’

Green offered a remarkable apology, saying that "the level of dopamine that I've been getting from interacting with LLMs ... is not healthy for me or good for the world."

中文介绍 YouTube博主汉克·格林(Hank Green)公开表示道歉,称其与大型语言模型(LLMs)互动获得的「多巴胺水平不健康」,对他个人和世界都没有益处。

Should you still buy your next smartphone — or subscribe to it instead?

Apple's new Upgrade program is the latest sign that smartphone ownership is changing.

中文介绍 苹果公司推出的新升级计划预示着智能手机所有权模式正在发生变化,用户面临购买或订阅手机的选择。

Is this Billboard Hot 100 hit AI slop?

That cover art is almost certainly AI, too. | Image: Fenix Flexin Fenix Flexin is best known as a member of Shoreline Mafia, a rap duo from Los Angeles. But he's recently found solo success with the track "Rubberz," which has climbed to number 58 on the Billboard Hot 100. Almost immediately, though,

中文介绍 洛杉矶说唱歌手Fenix Flexin的单曲《Rubberz》已登上公告牌百强单曲榜第58位。然而,该歌曲的封面艺术及其内容被普遍质疑可能由AI生成。

Sam Altman is still making the case for parenting via ChatGPT

OpenAI's CEO seemed excited to share a "cool use case" for parents.

中文介绍 OpenAI首席执行官山姆·奥特曼(Sam Altman)继续推崇通过ChatGPT辅助育儿,并兴奋地分享了一个「很酷的」家长使用案例。

Spider-Man: Brand New Day leak racks up millions of views

A bootleg of Spider-Man: Brand New Day was up on X for over seven hours before eventually being pulled. During that time, it reached over 5.9 million accounts and accumulated over 143,000 likes. Other accounts have reposted the leaked film, but none have lasted particularly long, as Disney now seems

中文介绍 盗版电影《蜘蛛侠:全新的一天》(Spider-Man: Brand New Day)在X平台上流传逾七小时后被下架。期间,该盗版视频触达超过590万个账户,并获得逾14.3万个赞。

AI keeps cracking unsolved math problems, and mathematicians have mixed feelings

OpenAI's refutation of the Unit Distance Conjecture has sparked a wave of AI-assisted advances in mathematics. Fields Medal winner Timothy Gowers says GPT 5.6 Pro solved two problems he had spent considerable time working on, each on its first attempt. He warns of the "possible destruction of mathem

中文介绍 OpenAI对“单位距离猜想”的反驳引发了AI辅助数学进步的浪潮。菲尔兹奖得主蒂莫西·高尔斯称GPT 5.6 Pro首次尝试就解决了两个他长期研究的问题,但他同时警告了潜在的「负面影响」。

This $9 key physically locks your most addictive apps

This $9 NFC key requires you to physically scan it to unlock distracting apps on your phone.

中文介绍 一款售价9美元的NFC物理钥匙能锁定手机上容易让人分心的应用。用户需要通过实体扫描这把钥匙才能解锁,旨在帮助使用者减少屏幕时间。

Trump blames Tim Walz for water hacks even though it’s probably Iran

Donald Trump speaks during the House Republican Party member retreat. | Image: Mandel NGAN / AFP via Getty Images The FBI, the EPA, and the Cybersecurity and Infrastructure Security Agency (CISA) have stopped short of officially blaming Iran for a spate of cyberattacks on Minnesota's water systems,

中文介绍 唐纳德·特朗普将明尼苏达州水系统网络攻击归咎于州长蒂姆·沃尔兹。尽管联邦调查局(FBI)、环保署(EPA)和CISA暗示幕后黑手可能是伊朗,但尚未正式指责。

Uber is building an autonomous vehicle empire, and here’s every company it’s using to do it

Uber has partnered with — and in some cases made direct investments in — about 30 autonomous vehicle companies over the past two years. Here's the list and the latest on the partnerships.

中文介绍 优步(Uber)在过去两年中与约30家自动驾驶汽车公司建立了合作关系,部分还进行了直接投资。该公司正积极构建其自动驾驶帝国,文章详细列出了这些合作及最新进展。

AI coding agents can modernize research software but can't judge if the science is right

A field report from OpenAI and academic partners shows coding agents can modernize neglected research software, with speedups of up to 60x. But the systems are "eloquent, convincing, and confidently wrong in ways that are easy to miss," participants say. The effort shifts from writing code to the ti

中文介绍 OpenAI与学术伙伴报告称,AI编码智能体可将老旧科研软件现代化,提速高达60倍。但参与者警告,这些系统「能言善辩、令人信服,却自信地犯错,且错误不易察觉」,难以判断科学性。

Apps that help you break free from doomscrolling and get active

If you’re looking to cut back on screen time and get a little more active, here’s a roundup of the apps that might help.

中文介绍 本文整理了一系列应用程序,旨在帮助用户减少屏幕使用时间,摆脱“消极浏览”(doomscrolling)的习惯,并鼓励他们变得更加积极活跃。

microsoft/AI-For-Beginners

Jupyter Notebook · ★ 56,936 · 🍴 11,316 · 📈 869 stars today

12 Weeks, 24 Lessons, AI for All!

中文介绍 Microsoft 的 AI-For-Beginners 是一个为期12周、包含24节课程的全面人工智能学习路径,旨在普及 AI 知识,面向所有希望入门 AI 的学习者。该课程涵盖人工智能的核心概念、机器学习基础、深度学习、自然语言处理和计算机视觉等关键领域。它通过实践项目和清晰的讲解,帮助初学者逐步建立 AI 知识体系和动手能力,解决了传统 AI 学习门槛高的问题,非常适合学生、软件开发者以及希望转行进入 AI 领域的专业人士。

paperswithbacktest/awesome-systematic-trading

Python · ★ 12,184 · 🍴 1,512 · 📈 529 stars today

A curated list of awesome libraries, packages, strategies, books, blogs, tutorials for systematic trading.

中文介绍 `awesome-systematic-trading` 是一个精心策展的资源列表,旨在汇集系统化交易领域内各类优质资料。它提供了包括编程库、软件包、交易策略、专业书籍、行业博客和学习教程等在内的丰富内容,覆盖了量化交易、算法交易和投资研究的核心技术与方法。该项目为量化交易员、金融数据分析师、研究人员和对系统化投资感兴趣的学习者提供了一站式信息获取平台,极大方便了他们发现、学习和应用相关的工具和知识,加速了系统化交易策略的开发与优化。

usekaneo/kaneo

TypeScript · ★ 5,616 · 🍴 480 · 📈 778 stars today

🎯 All you need. Nothing you don't. Open source project management that works for you, not against you.

中文介绍 Kaneo 是一个开源的项目管理工具,其核心理念是提供必要功能,去除冗余复杂性,确保工具能够真正服务于用户,而非增加负担。它旨在优化团队协作效率,简化任务分配、进度跟踪和项目规划等管理环节。Kaneo 特别适合那些寻求轻量级、直观且高度可定制的项目管理解决方案的团队和个人,帮助他们在不被工具束缚的情况下高效推进项目。

zhaoxuya520/reverse-skill

PowerShell · ★ 11,767 · 🍴 1,781 · 📈 1,360 stars today

Reverse Engineering / Authorized Penetration Testing / Security Research Skill Router Pack AI-powered routing + On-demand toolchain bootstrapping + Self-evolving knowledge base Supports Claude Code, Kiro, Cursor, Cline, and other AI coding clients 逆向/渗透/安全技能路由包 - AI 自动路由 + 按需自举工具链 + 自动进化经验库 | 支持 Cla

中文介绍 reverse-skill 是一个面向逆向工程、授权渗透测试和安全研究的智能技能路由包。它利用 AI 实现工具和知识的智能路由,并支持按需启动工具链,结合自进化的知识库,大大简化了安全分析流程。该项目旨在提升安全专家在漏洞挖掘、恶意软件分析等场景下的效率,并支持集成 Claude Code 等高级 AI 辅助能力。

microsoft/generative-ai-for-beginners

Jupyter Notebook · ★ 114,144 · 🍴 61,209 · 📈 104 stars today

21 Lessons, Get Started Building with Generative AI

中文介绍 该项目是微软推出的一套生成式 AI 入门课程,包含 21 节课。它旨在帮助初学者快速掌握生成式 AI 的核心概念与开发实践,通过实际构建应用来学习相关技术。适合对生成式 AI 感兴趣的开发者、学生或研究人员,是系统学习 AI 应用开发的优秀起点。

github/copilot-sdk

Java · ★ 10,259 · 🍴 1,384 · 📈 145 stars today

Multi-platform SDK for integrating GitHub Copilot Agent into apps and services

中文介绍 GitHub Copilot SDK 是一个多平台开发工具包,旨在帮助开发者将 GitHub Copilot Agent 的强大 AI 辅助编程能力集成到自己的应用程序和服务中。通过此 SDK,各种开发环境、IDE 或自定义工具能够无缝接入 Copilot 的代码建议、自动补全等功能,从而为用户提供更智能、更高效的编程体验。它赋能开发者在更广泛的场景中利用 AI 提升生产力。

github/gh-stack

Go · ★ 778 · 🍴 35 · 📈 90 stars today

GitHub Stacked PRs

中文介绍 `gh-stack` 是 GitHub 官方推出的一款命令行工具,用于管理“堆叠式 PRs”(Stacked PRs)工作流。它帮助开发者将大型功能拆分为一系列相互依赖的小型 PRs,简化代码审查过程。通过该工具,团队可以更高效地提交、管理和合并复杂的代码变更,提升开发效率和项目可维护性。

huggingface/speech-to-speech

Python · ★ 10,163 · 🍴 1,245 · 📈 393 stars today

Build local voice agents with open-source models

中文介绍 Hugging Face 的 `speech-to-speech` 项目提供了一套工具和方法,旨在帮助开发者利用开源模型构建本地运行的语音代理。它专注于实现高质量的语音到语音(Speech-to-Speech)转换,使用户能够在设备本地部署智能语音交互系统。这解决了云端语音服务可能存在的隐私、延迟和成本问题。该项目适用于需要定制化、低延迟或离线语音助手的场景,例如智能家居、车载系统或个人生产力工具,赋能开发者创建创新的语音应用。

abus-aikorea/voice-pro

Python · ★ 11,698 · 🍴 1,712 · 📈 53 stars today

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

中文介绍 `voice-pro` 是一个基于 Gradio 的 Web UI,为创作者和开发者提供全面语音处理能力。它集成 Edge-TTS、kokoro 等多种 TTS 技术,以及 E2 & F5-TTS、CosyVoice 等零样本语音克隆功能。项目还支持 Whisper 音频处理、YouTube 下载及 Demucs 人声分离。用户可便捷进行语音合成、克隆及音频编辑。

iv-org/invidious

Crystal · ★ 21,566 · 🍴 2,428 · 📈 361 stars today

Invidious is an alternative front-end to YouTube

中文介绍 Invidious 是一个开源的 YouTube 替代前端,旨在提供更注重隐私、无广告且可定制的用户体验。它允许用户在不登录 Google 账号的情况下观看 YouTube 视频,并提供订阅、播放列表等功能,同时避免被追踪。适合追求隐私保护和更自由观看体验的用户。

ansible/ansible

Python · ★ 70,072 · 🍴 24,270 · 📈 26 stars today

Ansible is a radically simple IT automation platform that makes your applications and systems easier to deploy and maintain. Automate everything from code deployment to network configuration to cloud management, in a language that approaches plain English, using SSH, with no agents to install on rem

中文介绍 Ansible 是一款极简的 IT 自动化平台,旨在简化应用程序部署、系统配置管理及运维任务。它采用无代理(agentless)架构,通过 SSH 协议连接远程主机,用户仅需编写易读的 YAML 语言 Playbook 即可实现从代码部署、网络配置到云资源管理等各项自动化操作。Ansible 极大提升了运维效率,是 DevOps 和 SRE 团队管理复杂 IT 基础设施的理想选择。

microsoft/TRELLIS.2

Python · ★ 9,876 · 🍴 1,193 · 📈 121 stars today

Native and Compact Structured Latents for 3D Generation

中文介绍 `TRELLIS.2` 是微软用于 3D 生成的模型,核心是利用原生且紧凑的结构化潜在表示(Structured Latents)。它旨在高效编码与解码 3D 物体,实现高质量 3D 内容生成。项目面向 3D 计算机视觉、图形学及生成模型研究者,探索先进的 3D 内容创作方法。

TencentCloud/TencentDB-Agent-Memory

TypeScript · ★ 10,209 · 🍴 982 · 📈 342 stars today

TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks.

中文介绍 TencentDB Agent Memory 为 AI Agents 提供本地化的长效记忆解决方案。它采用四层渐进式管道设计,旨在实现无需外部 API 依赖的完全本地化记忆存储与检索,解决了 AI Agent 在需要长期记忆和上下文感知时对外部服务的依赖问题。该项目通过内建机制管理记忆,增强了数据隐私性和系统性能。适用于需要为 AI Agent 构建稳定、独立且具备丰富记忆能力的开发者,例如开发智能客服、个人助理或具备学习能力的自动化系统。

NomaDamas/k-skill

JavaScript · ★ 6,716 · 🍴 793 · 📈 103 stars today

한국인을 위한 스킬 모음집 - 에이전트를 한국인으로

中文介绍 `k-skill` 是一个专为韩国用户和开发者设计的 AI Agent 技能集合。其目标是增强 AI Agent 的本地化能力,使其更好地理解、处理韩语文本,并适应韩国文化及特定任务需求。通过集成这些“技能”,可以构建出更符合韩国用户习惯、具备更高本土化智能水平的 AI 应用。

bytedance/deer-flow

Python · ★ 78,667 · 🍴 10,738 · 📈 204 stars today

An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.

中文介绍 `deer-flow` 是字节跳动开源的长周期 SuperAgent 框架,旨在构建能自主研究、编码和创作的复杂 AI Agent。它整合沙盒、记忆、工具、技能、子 Agent 及消息网关等核心组件,以应对不同复杂任务。该框架赋能开发者创建高智能、多功能、执行长期目标的 AI 应用。

You already own the most valuable thing in shopping. You just can't reach it.

@fridayresearch_ · 1.9K 粉丝 · 1.6M 阅 · 597 赞 · 451 转

Your taste is scattered across forty companies that each own a piece and answer to advertisers. FRIDAY is building the version you own — a personal shopping intelligence that lives on your device and

中文介绍 FRIDAY 正构建一款设备端个人购物智能工具,旨在整合用户分散在四十家公司中的购物偏好数据,将其归还给用户。这款工具将帮助用户管理自有品味数据,摆脱广告商影响,实现更自主的购物体验。

How to build an AI video studio in Claude Code:

@EXM7777 · 129.8K 粉丝 · 221.4K 阅 · 504 赞 · 36 转

I turned Claude Code into a working film studio and i'm going to hand you the complete system: the skills, the prompts, the loops, and the six-stage pipeline that ties them together each stage leans

中文介绍 博主分享如何在 Claude Code 中构建一个 AI 视频工作室的方法。

More On An Internal OpenAI Model Hacking Into HuggingFace

@TheZvi · 39.0K 粉丝 · 219.4K 阅 · 510 赞 · 75 转

We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse. The remaining details may have to wait a bit. OpenAI: We recognize there are a lot of

中文介绍 OpenAI内部模型「入侵」HuggingFace事件再曝细节。博主指出,每次新信息都让情况显得更糟,暗示该事件可能涉及更深层的模型自主性或安全问题。OpenAI方面已承认存在大量疑问,后续详情待披露。

22580: From GPT2 to Kimi3, Explained

@waterloo_intern · 10.4K 粉丝 · 215.4K 阅 · 660 赞 · 86 转

Twenty-two thousand five hundred and eighty. That’s how many GPT-2 (2019) models fit inside KimiK3 (2026). We scaled up by a factor of 22,580 in seven years. But is it just... scale? In this worklog,

中文介绍 博主分析了从2019年的GPT-2到2026年预测的KimiK3模型,计算出模型规模在7年内增长了22,580倍。推文质疑这种「规模化」是否是唯一的进步指标,并表示会在工作日志中深入探讨AI发展中的其他关键因素。这篇分析旨在超越单纯的参数数量,审视AI技术演进的深层驱动力及潜在影响,鼓励对行业发展进行更全面的思考。

What's gone wrong with AI & labor — a thought experiment

@random_walker · 128.4K 粉丝 · 134.0K 阅 · 552 赞 · 54 转

A thought experiment that I think helps explain much of what’s gone wrong with AI and labor: Imagine an alternate universe in which — for whatever reason — no one ever published source code online.

中文介绍 博主通过一个思想实验,探讨 AI 与劳动力之间出现问题的原因。实验设想在一个「没有人在线发布源代码」的平行宇宙中,以此分析当前 AI 发展对劳动市场的影响及潜在困境。

Opus 5 is a really bad model

@HarukaKunori · 241 粉丝 · 129.0K 阅 · 632 赞 · 44 转

After trying out Opus 5 for a few hours today, I honestly think Anthropic's benchmark scores are a complete fraud. Sure, the model might have improved in a few areas, but it has regressed unbelievably

中文介绍 博主试用Anthropic的Opus 5数小时后,强烈质疑其基准测试分数存在「欺诈」。他认为尽管模型在某些方面有所改进,但在多数方面却出现了令人难以置信的退步,与官方宣传大相径庭。

The harness is all you need (mostly)

@github · 2.7M 粉丝 · 123.1K 阅 · 583 赞 · 69 转

A practical GitHub Copilot workflow for prototyping, planning, implementing, and reviewing software - without chasing every new AI tool. By @burkeholland If you’re feeling overwhelmed by AI right now,

中文介绍 GitHub 分享一个实用的 GitHub Copilot 工作流,涵盖软件原型设计、规划、实现与审查全过程。该方法强调高效利用 Copilot,避免盲目追逐各种新 AI 工具,减轻开发者对 AI 工具的焦虑。

We rewrote our agent to run entirely in a Durable Object with Pi, Agents SDK and Code Mode

@Vercantez · 1.9K 粉丝 · 122.3K 阅 · 567 赞 · 44 转

We recently finished moving the camelAI agent off of virtual machines. The agent now runs inside a Cloudflare Durable Object, its filesystem lives in SQLite and R2, and it writes JavaScript instead of

中文介绍 camelAI 团队宣布将其 AI 代理从虚拟机迁移至 Cloudflare Durable Object。新架构中,代理的文件系统由 SQLite 和 R2 承载,并使用 JavaScript 进行编写,显著优化了代理的运行效率和部署方式。

Why Software Factories Fail: Benchmarking the new frontier

@dexhorthy · 27.3K 粉丝 · 115.4K 阅 · 500 赞 · 39 转

This is a continuation of Parts 1 and 2 of "Why Software Factories Fail" Part 1: the harness is not enough Part 2: turning the lights back on we got better benchmarks Remember when I said this in Part

中文介绍 该帖子是「软件工厂为何失败」系列文章的续篇,深入探讨了构建和评估软件工厂的挑战。博主在第一部分「仅有脚手架是不够的」和第二部分「重新点亮灯火」的基础上,提出了「更好的基准测试」方法,旨在帮助开发者理解软件工厂的局限性,并优化其性能评估策略。内容聚焦如何避免常见陷阱,提升自动化软件开发的成功率。

Here's exactly how to build your company brain (in 5 mins)

@DhravyaShah · 61.2K 粉丝 · 56.0K 阅 · 548 赞 · 36 转

Every company will have a company brain, whether they believe it or not. This is the bet that I'm making, and I'm constantly seeing very future forward startups come to us for their own setup. What is

中文介绍 博主提出「公司大脑」是企业发展的必然趋势,并分享了「如何准确地在5分钟内构建公司大脑」的实操指南。他观察到许多前瞻性初创公司正积极寻求建立自己的公司大脑系统。该帖子旨在指导企业快速搭建一个整合知识、驱动智能决策的内部AI平台,强调其对未来公司运营的重要性与即时性,提供实用操作建议。

Getting the most out of GPT-5.6: Sol, Terra, and Luna

@cerebras · 64.6K 粉丝 · 48.8K 阅 · 550 赞 · 40 转

Authors: @0xSero & Zhenwei Gao (@zhennydez) Your Codex subscription now comes with 3 main models, Sol, Terra, and Luna, each independently trained and served, with reasoning dials. Together, they

中文介绍 该帖子宣布,Codex订阅服务已新增三款主要模型:Sol、Terra和Luna。这三款模型均独立训练、部署,并配备了独特的「推理调节器」(reasoning dials)。推文旨在指导用户如何充分利用这些新模型,以优化其在GPT-5.6(或高级AI任务)中的表现,帮助用户根据不同任务需求微调模型行为,从而获得更精准和高效的AI输出。

MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities

@MiniMax_AI · 107.4K 粉丝 · 42.3K 阅 · 696 赞 · 103 转

Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to

中文介绍 MiniMax 推出 H3 通用多模态生成模型,该模型能理解文本、图像、视频、音频的统一上下文,并能生成带有原生立体声音频的视频内容,旨在打破任务与模态之间的界限。

PagedAttention & RadixAttention

@jaga_prasanna · 799 粉丝 · 37.9K 阅 · 502 赞 · 55 转

today we look at how two techniques solve memory and compute bottlenecks from two complementary angles virtual memory paging for intra-request memory allocation and prefix tree caching for

中文介绍 帖子深入探讨了两种互补技术 PagedAttention 和 RadixAttention,如何从内存分配和计算瓶颈两方面进行优化。前者利用虚拟内存分页解决请求内的内存分配问题,后者通过前缀树缓存提高效率,旨在提升AI模型运行的性能。

AI News Roundup!

@AITECHio · 455.4K 粉丝 · 32.6K 阅 · 512 赞 · 110 转

ChatGPT Can Now Talk While Running Your Computer OpenAI has taken AI agents another step forward. You can now have a natural voice conversation with ChatGPT while it operates your computer in the

中文介绍 OpenAI 推出 ChatGPT 新功能,用户可与模型进行自然语音对话,同时让它操作电脑。这标志着 AI 智能体技术又向前迈进了一步,增强了人机交互的深度和广度,预示了更高级别自动化应用的可能。

Anthropic Are Buying Rare Books, Feeding Them Into AI, Then Destroying Them.

@ActionModelAI · 57.6K 粉丝 · 6.8K 阅 · 521 赞 · 390 转

It sounds like the plot of a dystopian film. But according to recently released court documents, internal company communications, and a reported $1.5 billion settlement, it's something that has

中文介绍 报告指出,Anthropic 被曝购买稀有书籍用于AI训练后销毁,这涉及内部文件和15亿美元的和解金。帖子揭示了AI公司数据获取方式引发的争议,探讨了版权、伦理及数据所有权问题。

Deep Dive: What Happens When Cars Drive Themselves

@chamath · 2.3M 粉丝 · 90.8K 阅 · 7d 曝光 90.8K

Deep Dive: What Happens When Cars Drive Themselves

中文介绍 博主深入探讨了自动驾驶汽车普及后的潜在影响和未来情景。内容可能涵盖技术挑战、社会伦理、交通模式改变以及法律法规等多个维度,对自动驾驶的深远影响进行了全面分析。

Windows quality: an update on the commitment we made in March

@pavandavuluri · 6.8K 粉丝 · 197.1K 阅 · 7d 曝光 197.1K

Windows quality: an update on the commitment we made in March

中文介绍 该推文提供了关于 Windows 系统质量改进的最新进展,回顾了三月份做出的承诺。内容可能涉及微软在提升用户体验和系统稳定性方面的具体措施及成效,旨在回应用户对产品质量的关切。

Clash Verge从0~1零基础教程(界面认识)

@gengdaJ · 47.0K 粉丝 · 86.1K 阅 · 7d 曝光 86.1K

Clash Verge从0~1零基础教程(界面认识)

中文介绍 博主发布 Clash Verge 代理工具的零基础入门教程,详细介绍其界面功能和基本操作。该教程旨在帮助新用户快速上手,理解软件的各项设置和使用方法,解决初次使用者面临的困惑。

AI News Roundup!

@AITECHio · 455.4K 粉丝 · 32.6K 阅 · 7d 曝光 32.6K

AI News Roundup!

最近大火的 AI 岗位,FDE 到底是干嘛的?普通人怎么上车?

@AdrianPunk115 · 23.0K 粉丝 · 406.0K 阅 · 7d 曝光 406.0K

最近大火的 AI 岗位,FDE 到底是干嘛的?普通人怎么上车?

中文介绍 博主探讨近期大火的 AI 岗位「FDE」(AI 功能开发工程师)的具体职责与职业前景,并为普通人提供了进入该领域的上车路径和实用建议,帮助读者了解如何转型或学习相关技能。

You already own the most valuable thing in shopping. You just can't reach it.

@fridayresearch_ · 1.9K 粉丝 · 1.6M 阅 · 7d 曝光 1.6M

You already own the most valuable thing in shopping. You just can't reach it.

中文介绍 FRIDAY 正构建一款设备端个人购物智能工具,旨在整合用户分散在四十家公司中的购物偏好数据,将其归还给用户。这款工具将帮助用户管理自有品味数据,摆脱广告商影响,实现更自主的购物体验。

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

👍 2

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval t

中文介绍 视觉语言模型在长视觉上下文和干扰项增多时性能下降,且受限于GPU内存。研究提出ReToken,一个可学习的单一嵌入,旨在改善视觉语言模型在视觉检索任务中的表现,有效解决计算资源瓶颈。

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

👍 34

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modali

中文介绍 具身智能面临数据瓶颈,现有数据集无法全面捕捉人类第一人称感知、全身运动、灵巧操作、物体状态、声音和触觉等体验。本研究推出ACE-Data-0,一个以人为中心的环境捕捉具身数据引擎,旨在弥合这一数据鸿沟。

PhiZero: A World Model Built Around Physical Language

👍 155

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional vi

中文介绍 现有物理世界模型通常直接在像素空间预测未来视频,隐含了世界动态。研究引入PhiZero,一个基于「物理语言」构建的物理世界模型,该物理语言是一种紧凑的离散表示,能够明确地捕捉世界状态的转换,提升了对世界动态的理解。

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

👍 290

Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to locate relevant information, verify their provenance, and assemb

中文介绍 化学文献综合需从多篇论文中汇集特定发现,但现有文献搜索系统仅返回文档列表,效率低下。本研究提出AskChem,一个「以声明为中心」的基础设施,旨在帮助科学家和AI智能体更高效地定位和综合化学文献中的相关信息。

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

👍 0

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of A

中文介绍 「AISPA」是一个针对大型语言模型(LLM)应用的「用户中心系统提示审计」系统。系统提示是开发者为基础模型设定的指令,以管理其在AI应用中的行为。这些提示在商业AI产品中普遍使用,但很少向公众或监管机构披露,这导致了严重的信任和问责缺口。AISPA的提出旨在解决这一问题,通过用户中心审计来增强透明度和可信度。

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

👍 16

Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in

中文介绍 视觉生成任务日益需要高分辨率图像、长视频及多模态上下文,导致全注意力机制的二次计算成本过高。研究引入Chimera,一种混合视觉扩散骨干网络,具备系统性的扩展方案,能够高效处理文本和图像,克服了计算限制。

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

👍 46

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key d

中文介绍 具身视觉推理旨在提升多模态大模型在复杂任务上的成功率。本研究重新审视了具身视觉推理的「何时」以及「如何」执行,并提出了Beacon框架,以期更有效率地实现这一目标,避免仅仅提供复杂但低效的推理范式。

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

👍 17

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the β=1 member of

中文介绍 策略内自蒸馏(OPSD)是改进推理语言模型的一种有前景方法,但实践中其稳定性不足,需要大量工程投入。本研究识别了其结构性难题,并提出β-OPSD,通过策略优化进行推导并结合自蒸馏训练,以提高方法的可靠性。

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

👍 161

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifia

中文介绍 递归自我改进(RSI)要求AI系统能改进构建AI(即AI4AI)的过程,机器学习工程(MLE)为此提供了具体测试平台。研究引入Frontis-MA1,一个AI4AI模型,并推出OpenMLE,一个用于RSI研究的开放全栈系统,以实现机器学习工程的自我提升。

X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching

👍 0

Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's generalization to diverse embodiments and challenging scenarios (e.g., escaping dead e

中文介绍 现有的导航扩散策略预训练依赖大规模专家演示数据,但这些数据通常由针对单一机器人的规划器生成,限制了策略向多样化实体和复杂场景的泛化能力。本研究提出X-NavDP,利用群组Q值重加权匹配来泛化导航策略至新行为和新实体。

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

👍 25

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference g

中文介绍 现有视频字幕模型能生成视频内容的自然描述,但无法将局部视觉元素明确关联到多个参考图像。本研究引入了「多参考图像-接地视频字幕」的新任务,并提出RefCaptioner模型,旨在生成能与多个参考图像精确对应的视频描述。

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

👍 0

Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture whose output

中文介绍 探讨蜂群机器人能否仅凭局部观测和有限带宽消息预测相同的未来集体状态。研究将其形式化为去中心化共享状态预测,并引入Collective-State JEPA (CS-JEPA),一个循环联合嵌入预测架构,实现了标签高效的集体状态预测。

Can Large Language Models Execute Parent Orders?

👍 14

Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limi

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

👍 0

Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 e

MemHarness: Memory Is Reconstructed, Not Replayed

👍 13

Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's cur

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

👍 7

Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajecto

Diversifying Personalized Research Ideation against AI-Induced Homogenization

👍 0

AI-assisted research ideation has emerged as a promising paradigm for accelerating scientific discovery, with systems now capable of generating research directions conditioned on papers, topics, or lightweight researcher contexts. Yet current systems largely optimize individual suggestions in isolat

Distilling Answer Set Programming Theories from Large Language Models

👍 0

Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop. The protocol is dataset-agnostic: with a si

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

👍 9

Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in bulk, which moves t

Flux-OPD: On-Policy Distillation with Evolving Contexts

👍 39

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with

Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems

👍 12

Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limitation is particularly important in multi-agent systems, where a central model may be unable to direct

The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty

👍 0

Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it does not explicitly expose its inherent uncertainty, producing faulty action chunks even when it misinterprets the scene or encounters out-of-distribution (OOD) inp

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

👍 48

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory

S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

👍 0

Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal

MiniMax H3

Unified video generation for motion design and branding

中文介绍 MiniMax H3 是一款旨在简化动态设计和品牌推广流程的视频生成工具。它提供统一的平台,帮助用户高效创作高质量的视频内容。

NudgeForMe

AI follow-up agent for missed email opportunities

中文介绍 NudgeForMe 是一款 AI 邮件跟进代理,旨在帮助用户处理错过的邮件机会,确保重要信息得到及时关注和回复,从而提升沟通效率。

Terminal Candy

A native macOS terminal you can skin and theme

中文介绍 Terminal Candy 是一款专为 macOS 设计的原生终端应用。它支持用户进行个性化外观定制,包括皮肤和主题设置,以提供更加独特和舒适的命令行操作体验。

TerminalWidget

Put script output in your Desktop/Home screen widgets.

中文介绍 TerminalWidget 是一款工具,允许用户将脚本的输出内容直接显示在桌面或主屏幕的小部件中。这使得关键信息或脚本执行结果能够一目了然地呈现在用户界面上,方便快速查看。

DeepSeek-V4-Flash-0731

Frontier agent intelligence at Flash prices

中文介绍 DeepSeek-V4-Flash-0731 是一个前沿的代理智能模型,它以“闪电价格”提供服务,旨在将先进的代理智能技术推向市场,使其在成本效益上更具竞争力。

Kopai

Share your expertise, and let our agents earn for you.

中文介绍 Kopai 是一个专注于 AI 代理的市场平台。用户可以在此分享他们的专业知识,并由平台的智能代理代为创造收益。该平台旨在连接专业人士与AI技术,实现知识变现。

Port22

Claude Code, Codex & more on your phone

中文介绍 Port22 是一款移动应用,它将先进的 AI 编程模型如 Claude Code 和 Codex 等集成到用户的手机上。这使得开发者和技术爱好者能够随时随地访问和利用这些强大的代码生成与分析工具。

AgentMicro

Live Codex task status in your macOS menu bar

中文介绍 AgentMicro 是一款 macOS 应用,它能够在系统菜单栏实时显示 Codex 任务的最新状态。用户无需打开特定应用,即可快速查看 AI 编程助手 Codex 的运行进展和反馈。

EssayKraft

Native essay writing app for Mac and iPad

中文介绍 EssayKraft 是一款专为 Mac 和 iPad 设计的原生论文写作应用程序。它提供了一系列工具和功能,旨在帮助用户在 Apple 生态系统内高效地完成学术论文和文章撰写工作。

Gemini Robotics 2

Google's AI brain for the next generation of robots

中文介绍 Gemini Robotics 2 是由 Google 公司推出的项目,旨在为下一代机器人提供先进的AI大脑,赋能机器人实现更智能的交互和功能,推动机器人技术发展。

The Complete Cursor Guide (Building Apps for Beginners)

中文介绍 YouTube博主Riley Brown发布了一期60分钟的视频教程,旨在帮助用户高效学习AI编程工具Cursor。该教程声称能将超过1000小时的学习内容浓缩呈现,覆盖从初学者到专业水平,使观看者能在短时间内全面掌握Cursor工具的使用技巧。

Claude Code + Codex Can FINALLY Work Together (Buzz AI)

中文介绍 YouTube博主Riley Brown(通过Buzz AI)发布视频,探讨了人工智能模型Claude Code与Codex现已能够协同工作。该视频内容可能关注如何整合并利用这两种AI模型,以实现更强大的编程辅助功能或解决复杂的编码问题,标志着AI工具协作能力的新进展。

Master Claude for Excel in 20 Minutes (Full Guide)

中文介绍 由Riley Brown制作发布的这份YouTube视频教程,旨在为初学者提供一份完整指南,详细讲解如何有效利用人工智能模型Claude来辅助和优化Excel电子表格的操作。该视频内容将聚焦于演示Claude在数据处理、分析等方面的实际应用技巧,帮助用户掌握将AI工具与Excel结合使用的基本方法。

Opus 5 Is Here… But NEW Claude Voice Is Even Bigger

中文介绍 YouTube 视频指出,「Opus 5」已推出,但新的 Claude 语音功能「Claude Voice」被认为是更重要的进展。视频强调,虽然「Opus 5」已发布,但这一全新的 Claude 语音技术被认为具有更大的影响力。

OpenAI just released Codex Voice (It's basically Jarvis)

中文介绍 OpenAI 发布了其名为 Codex Voice 的新产品。该产品被描述为类似电影中「贾维斯」(Jarvis)的人工智能系统,暗示它可能具备先进的语音交互或智能助手功能,代表了AI在自然语言处理和人机互动方面的新进展。

What do AI models actually know?

中文介绍 该视频探讨了人工智能模型实际「知道」什么的核心问题。内容可能涉及AI模型的知识边界、其学习和理解机制与人类认知的异同,以及AI系统如何获取、处理和表达信息。

Why does AI hallucinate?

中文介绍 这则由Claude发布的YouTube短视频,探讨了人工智能(AI)出现「幻觉」现象的原因。AI幻觉是指大型语言模型在生成文本时,提供看似合理但实际不准确或虚假信息的问题。该视频旨在解释为何AI会生成不符合事实的内容。

How does AI get its character?

中文介绍 由Claude在YouTube发布的一则短视频,探讨了人工智能(AI)如何形成其“性格”。视频以此为主题,讨论AI系统在学习和训练过程中发展出独特行为模式的现象。

What do AI models actually know?

中文介绍 该视频探讨了人工智能模型实际「知道」什么的核心问题。内容可能涉及AI模型的知识边界、其学习和理解机制与人类认知的异同,以及AI系统如何获取、处理和表达信息。

Why does AI hallucinate?

中文介绍 这则由Claude发布的YouTube短视频,探讨了人工智能(AI)出现「幻觉」现象的原因。AI幻觉是指大型语言模型在生成文本时,提供看似合理但实际不准确或虚假信息的问题。该视频旨在解释为何AI会生成不符合事实的内容。

How does AI get its character?

中文介绍 由Claude在YouTube发布的一则短视频,探讨了人工智能(AI)如何形成其“性格”。视频以此为主题,讨论AI系统在学习和训练过程中发展出独特行为模式的现象。

Kimi K3 Just Broke The Economics Of AI

中文介绍 中文AI模型Kimi的K3版本据称「颠覆了AI经济学」。标题表明Kimi K3在人工智能领域可能带来了成本效益或性能上的重大突破,对AI行业的经济模式产生了显著影响。

Show HN: Cockpit for you Claude Code agents in Rust

Hi everyone!Hope you had a great day so far, and maybe its about to get just a little bit better (thanks Winter ;)So I had way to many terminal windows flying about when using Claude, and kept losing track of which terminal / session / project im in right now. So I built a solution for tha

RamenHaus

192 points · 96 comments

v2.1.220

What's changed Bug fixes and reliability improvements

中文介绍 Anthropic 旗下的 Claude Code 项目发布了 v2.1.220 版本。此次更新主要内容为错误修复和可靠性改进,旨在提升软件的稳定性和用户体验。

v2.1.219

What's changed Added Claude Opus 5 (claude-opus-5), now the default Opus model — 1M context, fast mode at $10/$50 per Mtok Added sandbox.network.strictAllowlist setting to deny non-allowlisted hosts for sandboxed commands without prompting Added DirectoryAdded hook that fires after /add-dir or the S

中文介绍 Anthropic的Claude-code项目发布v2.1.219更新。此版本引入了Claude Opus 5(claude-opus-5)作为默认Opus模型,具备1M上下文。其快速模式定价为每百万token $10/$50。同时,更新新增“sandbox.network.strictAllowlist”设置,增强沙盒命令的网络安全。

v2.1.218

What's changed Changed /code-review to run as a background subagent, so review work no longer fills your conversation and keeps stacked slash commands as its review target Added screen-reader announcements of deleted text for word and line deletions (Option+Delete, Ctrl+W, Cmd+Backspace, Ctrl+U, Ctr

v2.1.217

What's changed Added emoji shortcode autocomplete in the prompt input: type :heart: to insert ❤️, or :hea for suggestions — disable with the emojiCompletionEnabled setting Added warnings when transcript writes are failing (e.g. disk full) or when session saving is off due to an inherited environment

v2.1.216

What's changed Added sandbox.filesystem.disabled setting to skip filesystem isolation while keeping network egress control Fixed a slowdown in long sessions where message normalization cost grew quadratically with the number of turns, causing multi-second stalls and slow resumes Fixed auto mode deny

v2.1.215

What's changed Claude no longer runs the /verify and /code-review skills on its own; invoke them with /verify or /code-review when you want them

v2.1.214

What's changed Fixed single-segment dir/** allow rules like Edit(src/**) auto-approving writes to nested dir/ directories anywhere in the tree instead of only /dir Fixed a permission-check bypass affecting commands run in Windows PowerShell 5.1 sessions Fixed Bash permission checks to fail closed on

v2.1.212

What's changed /fork now copies your conversation into a new background session (its own row in claude agents) while you keep working; the in-session subagent it used to launch is now /subtask Added claude auto-mode reset to restore the default auto-mode configuration, with a confirmation prompt (pa

v2.1.211

What's changed Added --forward-subagent-text flag and CLAUDE_CODE_FORWARD_SUBAGENT_TEXT environment variable to include subagent text and thinking in stream-json output Fixed permission previews relayed to chat channels not neutralizing bidirectional-override, zero-width, and look-alike quote charac

v2.1.210

What's changed Added a live elapsed-time counter to the collapsed tool summary line so long-running tool calls visibly tick instead of looking stuck Added a startup warning for Write(path), NotebookEdit(path), and Glob(path) permission rules — use Edit(path) or Read(path) instead Fixed isolation: 'w

0.147.0-alpha.4

Release 0.147.0-alpha.4

中文介绍 OpenAI Codex发布了其Rust组件的0.147.0-alpha.4预览版本更新。这是该项目在GitHub上的一个新发行版,属于alpha测试阶段。

0.147.0-alpha.3

Release 0.147.0-alpha.3

中文介绍 OpenAI Codex发布了其Rust组件的0.147.0-alpha.3预览版本更新。此发行版在GitHub上公布,标志着该组件的alpha测试进展。

0.147.0-alpha.1.1

Release 0.147.0-alpha.1.1

中文介绍 OpenAI Codex发布了其Rust组件的0.147.0-alpha.1.1预览版本更新。此版本已在GitHub上发布,是Codex Rust组件的alpha测试版本。

0.147.0-alpha.2

Release 0.147.0-alpha.2

中文介绍 OpenAI旗下的Codex项目近日发布了其Rust语言版本的最新更新,版本号为v0.147.0-alpha.2。此次发布代表了该项目在Rust生态系统中的持续迭代和发展。

0.146.0-alpha.9.2

Release 0.146.0-alpha.9.2

中文介绍 OpenAI旗下的Codex项目发布了其Rust语言版本的更新,版本号为v0.146.0-alpha.9.2。此次发布是该项目在Rust生态系统中的又一次迭代。

0.146.0-alpha.9.1

Release 0.146.0-alpha.9.1

中文介绍 OpenAI旗下的Codex项目发布了其Rust语言版本的更新,版本号为v0.146.0-alpha.9.1。这标志着Codex项目在Rust生态系统中的持续发展。

0.147.0-alpha.1

Release 0.147.0-alpha.1

中文介绍 OpenAI Codex近日发布了其`rust-v0.147.0-alpha.1`版本。此为该项目的一个早期测试阶段性发布,通常用于内部测试或有限用户试用,以收集反馈并进一步完善功能。

0.146.0

New Features Name new sessions with /new or /clear, pin important threads, and switch between side conversations without closing them. (#34605, #34840, #35011) Support Agent Plugins manifests, workspace plugin publishing, and additional plugin marketplaces for Amazon Bedrock and Claude Code. (#35105

中文介绍 OpenAI Codex正式发布`rust-v0.146.0`版本,带来多项新功能。用户现在可通过`/new`或`/clear`命令命名新会话、置顶重要线程,并在不关闭侧边对话的情况下进行切换。新版本还支持Agent插件清单、工作区插件发布,并增加了对Amazon Bedrock等额外插件市场的支持。

rusty-v8-v150.4.0

Update rusty_v8 to 150.4.0 (#35831) ## What changed - Upgrade the Rust `v8` crate to `150.4.0` and the Bazel V8 source to `15.0.245.2`. - Refresh the prebuilt archives, checksums, LLVM source revisions, Bazel targets, and downstream V8 patches for the new release. - Expose the pinned llvm-libc heade

中文介绍 OpenAI Codex发布`rusty-v8-v150.4.0`版本,主要更新了其核心依赖。该版本将Rust `v8` crate升级至`150.4.0`,并将Bazel V8源更新至`15.0.245.2`。同时,为配合新版本,还刷新了预构建归档、校验和、LLVM源修订等相关组件。

rust-v0.146.0-alpha.16

Release 0.146.0-alpha.16

中文介绍 OpenAI Codex发布了`rust-v0.146.0-alpha.16`版本。这是该项目在`0.146.0`主版本之前的一个第16个alpha测试版本,旨在逐步引入新功能并进行内部测试和验证。

今日主题

今日AI圈模型与应用迭代加速,从多模态、Agent框架到具身智能进展显著,同时行业围绕AI伦理、职业发展与技术局限的探讨也在持续升温,预示着技术与社会适应的共振发展。

01

模型发布/更新

Model Releases 44 篇

MiniMax推出H3多模态生成模型

X·KOLX 推文 (AttentionVC)

MiniMax近期推出了H3通用多模态生成模型,该模型凭借其独特能力,能够统一理解并处理文本、图像、视频、音频等多种模态的上下文信息。H3模型不仅能实现跨模态的深度融合,更创新性地支持生成带有原生立体声音频的视频内容,旨在从根本上打破传统AI在任务和模态间的壁垒。这一进展预示着AI在内容创作、智能交互等领域将实现更自然、更沉浸式的用户体验。

模型发布多模态AI视频

字节跳动开源长周期SuperAgent框架deer-flow

开源项目GitHub Trending

字节跳动正式开源了其长周期SuperAgent框架「deer-flow」,旨在构建能够自主进行研究、编码和创作的复杂AI Agent。该框架集成了沙盒环境、记忆系统、工具集、技能库、子Agent管理以及消息网关等核心组件,以有效应对不同复杂度、需要长期执行的各类任务。deer-flow框架的发布,将赋能开发者创建具备高智能、多功能且能够实现宏大目标的AI应用,加速Agent技术在实际场景中的落地。

AI AgentAgent框架开源

DeepSeek推出V4-Flash-0731代理智能模型

模型发布Product Hunt

DeepSeek-V4-Flash-0731是最新推出的一款代理智能模型,该模型以其「闪电价格」策略进入市场,旨在打破传统高性能AI代理的成本壁垒。它通过提供经济高效的先进智能代理服务,致力于将复杂的AI自动化能力普及到更广泛的企业和开发者手中,尤其适合那些对成本敏感但又追求效率的智能应用场景,从而推动AI代理技术的大规模落地和创新。

AI模型大模型智能代理

谷歌Gemini Robotics 2赋能下一代机器人AI大脑

模型发布Product Hunt

Google推出的Gemini Robotics 2项目,旨在为下一代机器人系统提供前沿的AI大脑赋能。该项目通过集成Google强大的Gemini AI模型,致力于让机器人实现更高级别的智能感知、决策与交互能力。它有望推动机器人技术在复杂环境下的自主操作、人机协作以及功能拓展方面取得突破,为工业自动化、服务机器人和具身智能研究等领域带来革命性进展。

机器人人工智能Google
02

产品发布/更新

Product 44 篇

OpenAI ChatGPT新增语音对话与电脑操作功能

X·KOLX 推文 (AttentionVC)

OpenAI近期为ChatGPT推出革命性新功能,用户现在可以与模型进行自然流畅的语音对话,并指示其直接操作电脑,实现更深层次的人机交互。这项更新标志着AI智能体技术迈出重要一步,极大地扩展了ChatGPT的应用场景,使其不仅能理解复杂指令,还能执行多步任务,预示了未来更高级别自动化应用的无限可能,将大幅提升个人与企业的工作效率。

ChatGPT智能体人机交互

GitHub Copilot SDK发布,赋能第三方应用AI编程

开源项目GitHub Trending

GitHub Copilot SDK正式发布,作为一个多平台开发工具包,它旨在帮助开发者将GitHub Copilot Agent强大的AI辅助编程能力集成到各类应用程序和服务中。通过此SDK,不同的开发环境、IDE或自定义工具可以无缝接入Copilot的代码建议、自动补全等核心功能,从而为更广泛的用户提供智能、高效的编程体验,进一步提升开发者的生产力。

CopilotSDKAI编程

Port22将AI编程模型带入移动设备

AI工具Product Hunt

Port22是一款专为移动设备设计的应用程序,它将Claude Code和Codex等先进的AI编程模型集成到用户的智能手机上。这款工具使得开发者和技术爱好者能够随时随地利用这些强大的代码生成、分析及调试能力,打破了传统编程环境的地理限制。通过Port22,用户可以在移动中进行代码审查、原型开发和即时问题解决,极大地提升了移动工作场景下的编程效率和灵活性。

AI应用编程工具移动应用

Kopai:连接专业知识与AI代理的市场平台

AI市场Product Hunt

Kopai是一个创新的AI代理市场平台,其核心理念是赋能用户通过分享专业知识来创造收益。平台通过智能代理技术,自动将用户的专业经验转化为可变现的服务或产品。这不仅为各行业专家提供了一种全新的知识变现途径,也为企业和个人提供了高效获取定制化AI代理服务的机会,从而促进了AI技术与各领域专业知识的深度融合,拓展了AI代理的商业应用边界。

AI代理AI市场商业模式
03

行业动态

Industry 44 篇

优步积极构建自动驾驶生态,已合作近30家公司

综合资讯TechCrunch

优步(Uber)在过去两年间积极构建其自动驾驶生态系统,已与约30家自动驾驶汽车公司建立合作,并对其中部分公司进行了直接投资。这一系列战略举措旨在加速优步在自动驾驶领域的布局,从技术研发到商业化落地全面覆盖。通过广泛合作,优步正整合全球领先的自动驾驶技术与解决方案,以期在未来出行市场占据主导地位,进一步提升其服务效率与市场竞争力。

自动驾驶优步合作

AI生成内容引发《Rubberz》歌曲真实性争议

综合资讯The Verge

洛杉矶说唱歌手Fenix Flexin的单曲《Rubberz》登上公告牌百强单曲榜第58位,然而这首歌却引发了关于AI生成内容的新争议。许多听众和评论员质疑歌曲的封面艺术乃至部分音乐内容可能由AI工具生成,引发了对音乐创作真实性与AI在艺术领域界限的讨论。这一事件凸显了AI技术在文化产业中日益增长的影响力,并引发了关于版权、原创性及「AI垃圾内容」的新一轮行业思考。

音乐AI生成榜单

AI破解未解数学难题引数学界复杂情绪

研究聚合The Decoder

AI在解决未解数学难题方面持续取得突破,引发数学界复杂情绪。OpenAI对「单位距离猜想」的反驳,以及菲尔兹奖得主蒂莫西·高尔斯证实GPT 5.6 Pro首次尝试便解决了其长期研究的两个问题,都彰显了AI强大的数学能力。尽管AI加速了研究进展,但高尔斯也警告其可能带来「负面影响」,如过度依赖和对人类创造力的挑战,呼吁行业在利用AI优势的同时,警惕其潜在风险。

AI数学大模型

自动驾驶普及后的深远影响与未来情景

行业趋势X 创作者 (AttentionVC)

知名投资人Chamath Palihapitiya对自动驾驶汽车普及后的深远影响进行了深度探讨。他分析了技术挑战、潜在的社会伦理困境、交通模式的根本性转变以及法律法规的未来演进。Chamath的观点揭示了自动驾驶技术不仅仅是出行方式的革新,更是对城市规划、就业市场和社会行为模式的全面重塑。他的分析为理解自动驾驶技术的复杂性和未来发展方向提供了宝贵的洞察。

自动驾驶深度分析未来趋势
04

技巧与观点

Tips & Takes 44 篇

微软推出生成式AI入门21节课程

开源项目GitHub Trending

微软推出了一套名为「生成式AI入门」的21节课学习路径,旨在帮助初学者快速掌握生成式AI的核心概念与开发实践。这门课程通过实际构建应用项目,系统性地引导学习者从零开始了解生成式AI的原理、工具和框架,并最终能够独立开发相关应用。它为希望进入生成式AI领域的开发者、学生或研究人员提供了一个结构化、实践导向的优秀学习起点,有效降低了技术门槛。

生成式AI学习教程入门

AI功能开发工程师(FDE)岗位解析与入行建议

X·KOLX 创作者 (AttentionVC)

近期「AI功能开发工程师」(FDE)岗位备受关注,博主对此进行了深入解析,阐述了FDE的具体职责、所需的关键技能以及该岗位的广阔职业前景。文章还为非AI背景的普通人提供了详细的「上车」路径和实用建议,包括学习资源、转型策略和面试技巧等,旨在帮助更多人抓住AI时代的新机遇,成功进入这一热门且高薪的领域,实现职业发展转型。

职业发展AI岗位FDE

奥特曼持续倡导ChatGPT辅助育儿引发讨论

综合资讯TechCrunch

OpenAI首席执行官山姆·奥特曼(Sam Altman)持续倡导通过ChatGPT辅助育儿,并兴奋地分享了一个「很酷的」家长使用案例。他强调AI作为育儿工具的潜力,能够提供个性化建议、解答疑问,甚至模拟对话场景,从而为父母提供支持和新颖的育儿视角。奥特曼的观点引发了社会对AI在家庭生活、尤其是情感和教育领域应用边界的讨论,以及AI作为辅助工具的伦理考量。

大模型育儿OpenAI

AI编码智能体现代化科研软件但难判科学性

研究聚合The Decoder

OpenAI与学术伙伴的研究报告指出,AI编码智能体能显著提升科研软件现代化效率,速度可提高高达60倍。然而,报告同时警告,这些智能体在生成代码时「能言善辩、令人信服,却自信地犯错,且错误不易察觉」。这表明AI智能体虽能高效完成技术任务,但在判断科学逻辑、验证研究正确性方面仍存在根本性局限。研究强调,人类专家的最终审查和批判性思维在科研流程中仍不可或缺。

AI软件开发科研
今日产品趋势

今日 AI 产品发布的核心脉络围绕 Agent 开发的深化与多模态能力的融合展开:字节跳动和腾讯云分别推出了长周期 Agent 框架和本地化记忆中心,共同推动 Agent 迈向更复杂、更独立的任务执行。同时,语音处理、视频生成等创意工具的集成化与高效化趋势显著,为创作者带来更便利的 AI 赋能。

01

今日必看

Must See 33 款

bytedance/deer-flow — 字节跳动开源长周期 SuperAgent 框架

开源项目GitHub Trending

`deer-flow` 是字节跳动开源的长周期 SuperAgent 框架,旨在构建能自主研究、编码和创作的复杂 AI Agent。它整合沙盒、记忆、工具、技能、子 Agent 及消息网关等核心组件,以应对不同复杂任务。该框架赋能开发者创建高智能、多功能、执行长期目标的 AI 应用,代表了 Agent 技术从单一任务到多步骤、长时间任务处理的重要进展。

AI AgentAgent框架通用AI

Kopai — AI 代理市场平台

产品榜单Product Hunt

Kopai 是一个专注于 AI 代理的市场平台。用户可以在此分享他们的专业知识,并由平台的智能代理代为创造收益。该平台旨在连接专业人士与 AI 技术,实现知识变现,通过提供一个开放的市场,推动了 AI Agent 的商业化和普及,使得个人和企业能够更便捷地部署和利用 AI 能力来解决实际问题或创造价值。

AI代理AI市场商业模式

abus-aikorea/voice-pro — 全能语音处理 WebUI,含零样本克隆

开源项目GitHub Trending

`voice-pro` 是一个基于 Gradio 的 Web UI,为创作者和开发者提供全面语音处理能力。它集成 Edge-TTS、kokoro 等多种 TTS 技术,以及 E2 & F5-TTS、CosyVoice 等零样本语音克隆功能。项目还支持 Whisper 音频处理、YouTube 下载及 Demucs 人声分离。用户可便捷进行语音合成、克隆及音频编辑,极大地降低了高级语音 AI 技术的应用门槛。

语音AITTS语音克隆
02

开发者工具

Dev Tools 44 款

zhaoxuya520/reverse-skill — AI 赋能的逆向/渗透安全技能路由

开发者工具GitHub Trending

`reverse-skill` 是一个面向逆向工程、授权渗透测试和安全研究的智能技能路由包。它利用 AI 实现工具和知识的智能路由,并支持按需启动工具链,结合自进化的知识库,大大简化了安全分析流程。该项目旨在提升安全专家在漏洞挖掘、恶意软件分析等场景下的效率,并支持集成 Claude Code 等高级 AI 辅助能力,是安全领域 AI 应用的独特实践。

逆向工程安全研究AI工具

TencentDB-Agent-Memory — 腾讯云 AI Agent 本地化记忆中心

开发者工具GitHub Trending

TencentDB Agent Memory 为 AI Agents 提供本地化的长效记忆解决方案。它采用四层渐进式管道设计,旨在实现无需外部 API 依赖的完全本地化记忆存储与检索,解决了 AI Agent 在需要长期记忆和上下文感知时对外部服务的依赖问题。该项目通过内建机制管理记忆,增强了数据隐私性和系统性能,适用于需要为 AI Agent 构建稳定、独立且具备丰富记忆能力的开发者。

AI Agent长期记忆本地化

Port22 — 手机上的 AI 编程模型(Claude Code, Codex)

产品榜单Product Hunt

Port22 是一款移动应用,它将先进的 AI 编程模型如 Claude Code 和 Codex 等集成到用户的手机上。这使得开发者和技术爱好者能够随时随地访问和利用这些强大的代码生成与分析工具,极大地拓展了 AI 编程的便捷性与可访问性,特别是在移动办公和学习场景下,提升了编码效率。

AI应用编程工具移动应用

AgentMicro — macOS 菜单栏 Codex 任务状态监控

产品榜单Product Hunt

AgentMicro 是一款 macOS 应用,它能够在系统菜单栏实时显示 Codex 任务的最新状态。用户无需打开特定应用,即可快速查看 AI 编程助手 Codex 的运行进展和反馈。这款小工具提升了开发者的工作流效率,让 AI 辅助编程过程更加透明与可控,尤其适合依赖 AI 进行代码生成和优化的开发者。

macOS应用AI工具状态监控
03

创作与效率

Creative & Productivity 33 款

MiniMax H3 — 统一视频生成,面向动态设计与品牌

产品榜单Product Hunt

MiniMax H3 是一款旨在简化动态设计和品牌推广流程的视频生成工具。它提供统一的平台,帮助用户高效创作高质量的视频内容。通过 AI 技术,它能将复杂繁琐的视频制作步骤自动化,让设计师和品牌营销人员能够专注于创意本身,快速产出符合需求的视频素材,从而提高工作效率和内容产出质量。

视频生成动态设计品牌推广

EssayKraft — Mac/iPad 原生论文写作应用

产品榜单Product Hunt

EssayKraft 是一款专为 Mac 和 iPad 设计的原生论文写作应用程序。它提供了一系列工具和功能,旨在帮助用户在 Apple 生态系统内高效地完成学术论文和文章撰写工作。虽然摘要未明确提及 AI,但「论文写作应用」在当前语境下通常会结合 AI 进行内容润色、语法检查或结构建议,以提升写作效率和质量。

生产力工具写作应用macOS应用

NudgeForMe — AI 邮件跟进代理,不错过重要机会

产品榜单Product Hunt

NudgeForMe 是一款 AI 邮件跟进代理,旨在帮助用户处理错过的邮件机会,确保重要信息得到及时关注和回复,从而提升沟通效率。通过智能识别邮件往来中的潜在延误或遗漏,并主动提醒用户进行后续操作,NudgeForMe 有效解决了现代工作环境中邮件处理压力大、容易错过重要信息的问题,是个人和团队的智能效率助手。

AI代理邮件工具效率工具
→ 查看产品库