→ 模型会判断你的提示词是否需要真实世界的上下文。如果你要求"一张新 iPhone 18 的照片",它会先运行一次网络搜索,以了解它实…
分享链接 · https://openimages.ajiang.me/p/2074868654101901657/

color_palette · high-contrastcolor_palette · neoncomposition · close-upcomposition · multi-panellighting · neon-lightlighting · rim-light
提示词(原文,逐字保留)
→ The model decides whether your prompt needs real-world context. If you ask for "a photo of the new iPhone 18," it runs a web search first to know what it actually looks like.
→ It writes and executes code to solve accuracy problems it can't fix through generation alone. Mathematical plots, QR codes, precise diagrams — it renders them with Python, then uses that output as conditioning for the final image.
→ After generating, it evaluates its own work. Needs a fix? It does a targeted local edit. Needs a redo? It regenerates entirely. This self-refinement loop wasn't hand-coded — it emerged from reinforcement learning because better images earned higher rewards.
The result: more test-time compute = measurably better images. That's the same scaling principle that made o1 and DeepSeek-R1 work for reasoning, now applied to visual generation.
The benchmarks back it up. On Arena's image leaderboard (7,715 human preference votes), Muse Image ranks second in text-to-image, single-image editing, and multi-image editing. Only GPT Image 2 scores higher. It's the only top-tier model using agentic tools at inference time.
Muse Video launched alongside it — currently third on Arena's video generation board. Both share infrastructure with Muse Spark, MSL's language model, enabling joint planning on complex requests. A Muse Image output can become an interactive website or a video game. That's not an image generator. That's a media production pipeline.
The catch: no API yet. Muse Image lives inside Meta AI, Instagram Stories, and WhatsApp in select countries. Meta says a developer API is coming, but Muse Spark has been "coming soon" since April with no public access. Plan accordingly.
What excites me most isn't the benchmark scores. It's the paradigm shift: we went from "models that generate" to "models that reason about what to generate, generate, then fix their own mistakes." This is what agentic AI actually looks like in practice — not a chatbot calling tools, but the model itself treating tool use as part of its thinking process.
If this approach scales to video, 3D, and audio — and there's no reason it won't — every creative tool pipeline gets rewritten.
What's your take: is "agentic generation" the real next step for creative AI, or just a clever use of test-time compute that could be replicated by chaining existing tools together?
#MetaMuse
中文译文
→ 模型会判断你的提示词是否需要真实世界的上下文。如果你要求"一张新 iPhone 18 的照片",它会先运行一次网络搜索,以了解它实际的样子。
→ 它会编写并执行代码,以解决仅靠生成无法解决的准确性问题。数学图表、二维码、精确图示——它用 Python 渲染这些内容,然后将输出作为最终图像的条件。
→ 生成完成后,它会评估自己的作品。需要修复?它会进行有针对性的局部编辑。需要重做?它会完全重新生成。这种自我优化循环并非手工编写——它源自强化学习,因为更好的图像能获得更高的奖励。
结果是:更多的测试时算力 = 可衡量的更好图像。这与让 o1 和 DeepSeek-R1 在推理上有效的扩展原则相同,现在被应用于视觉生成。
基准测试也证明了这一点。在 Arena 的图像排行榜(7,715 次人类偏好投票)上,Muse Image 在文本生成图像、单图编辑和多图编辑中均排名第二。只有 GPT Image 2 的得分更高。它是唯一在推理时使用智能体工具的顶级模型。
Muse Video 与之同步发布——目前在 Arena 的视频生成排行榜上位列第三。两者都与 MSL 的语言模型 Muse Spark 共享基础设施,能够对复杂请求进行联合规划。Muse Image 的输出可以变成一个交互式网站或一款电子游戏。那不是图像生成器,那是媒体生产流水线。
不足之处是:目前还没有 API。Muse Image 部署在 Meta AI、Instagram Stories 和部分国家的 WhatsApp 中。Meta 表示开发者 API 即将推出,但 Muse Spark 自四月以来一直"即将上线",却没有任何公开访问。请据此规划。
最让我兴奋的不是基准测试分数,而是范式的转变:我们从"会生成的模型"走向了"对要生成的内容进行推理、生成、然后修正自身错误的模型"。这才是智能体 AI 在实践中的真实面貌——不是调用工具的聊天机器人,而是模型本身将工具使用视为其思维过程的一部分。
如果这种方法能够扩展到视频、3D 和音频——没有理由不能——那么每一条创意工具流水线都将被重写。
你的看法是什么:"智能体生成"是创意 AI 真正的下一步,还是仅仅是对测试时算力的巧妙利用,可以通过将现有工具串联起来来复制?
#MetaMuse
来源与署名
原文由 @Huintellimance 发布在 X:查看原推文。 本页逐字保留原文并提供机器翻译的中文解读;版权归原作者所有。
同工具的更多提示词
- 凌乱的构图,杂乱的背景,卡通插画,深色浑浊的茶,普通的杯子,模糊的茶包细节,过多的道具,喧闹的排版,随机多彩的背景,低细节的液体,浓稠的咖啡状泡沫,混乱的涂鸦…@ou_zhen599
- 蒙得维的亚,就像一部被尘封在抽屉里的80年代电影海报 世界上最长的滨海大道(rambla),在这里却像一个布景——那银灰色的光线让你彻底看呆 你想看到哪座城市…@MiMundoConIA
- https://t.co/0qRsl8vvEU 电影感的时尚摄影,场景设在纽约市一座高楼的屋顶上,天空为柔和漫射的阴天,密集的曼哈顿天际线向远处延伸。一位神情…@iamaiistudio
- “AI 工作流速查表,整洁的图标,编号步骤,蓝白配色,网页文章标题图”。很想看看你的版本。https://t.co/c1XCaHu8Yk@getimg_ai
- 一张蓝图风格的未来主义跑车技术图纸。包含跑车的前视图、侧视图和后视图的线条画、爆炸零件草图、零件装配图以及拆解组件的结构图。使用大量线条和测量数值标注每个零件…@marmaduke091
- 不要改变面部特征。照片宽高比 3:4。 一张具有“温馨时刻”美学的竖版家居风格照片,使用前置摄像头/自拍摄像头近距离拍摄。一位皮肤白皙、长直发的年轻女子站在一…@oggii_0