<?xml version='1.0' encoding='utf-8'?>
<rss version="2.0"><channel><title>Starsynced AI Wire — Frontier Model Releases</title><link>https://vps.starsynced.net/</link><description>Latest frontier model releases, refreshed every five minutes.</description><lastBuildDate>Sat, 12 Sep 2026 18:11:54 +0000</lastBuildDate><ttl>5</ttl><item><title>Perplexity trusts GPT-6 Astra with end-to-end systems</title><link>https://openai.com/index/perplexity-improving-accuracy-with-astra</link><guid isPermaLink="true">https://openai.com/index/perplexity-improving-accuracy-with-astra</guid><description>Perplexity uses Astra to write communications, change software, and monitor production systems, and checks in much less frequently than with earlier models.</description><author>OpenAI</author><pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate></item><item><title>Devin GPT-6 Astra Self-Testing Moves Review to Evidence</title><link>https://pastagi.com/use-cases/devin-gpt-6-astra-self-testing-review-bottleneck/</link><guid isPermaLink="true">https://pastagi.com/use-cases/devin-gpt-6-astra-self-testing-review-bottleneck/</guid><description>Devin GPT-6 Astra self-testing shifts code review from reading diffs to auditing evidence. Here is the pattern, its failure modes, and audit criteria.</description><author>Pastagi</author><pubDate>Sat, 12 Sep 2026 15:07:14 +0000</pubDate></item><item><title>One Capital Letter Was Silently Breaking My AI Support Bot, and It Wasn't in the New Model</title><link>https://towardsdatascience.com/one-capital-letter-was-silently-breaking-my-ai-support-bot-and-it-wasnt-in-the-new-model/</link><guid isPermaLink="true">https://towardsdatascience.com/one-capital-letter-was-silently-breaking-my-ai-support-bot-and-it-wasnt-in-the-new-model/</guid><description>A real Weave project that regression-tests three OpenAI models against the exact reply format your app depends on.</description><author>Towardsdatascience</author><pubDate>Sat, 12 Sep 2026 14:00:02 +0000</pubDate></item><item><title>From Hacks to Bioweapons, Claude Misuse Is Now Everywhere</title><link>https://www.wired.com/story/security-news-this-week-from-hacks-to-bioweapons-claude-misuse-is-now-everywhere/</link><guid isPermaLink="true">https://www.wired.com/story/security-news-this-week-from-hacks-to-bioweapons-claude-misuse-is-now-everywhere/</guid><description>Plus: The US disrupts the internet’s biggest black market, a Conti ransomware hacker gets prison time, Meta fails to stop AI-generated videos of child abuse.</description><author>Wired</author><pubDate>Sat, 12 Sep 2026 10:30:00 +0000</pubDate></item><item><title>[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale</title><link>https://www.latent.space/p/ainews-deepseek-v41-flash-763b-p8b</link><guid isPermaLink="true">https://www.latent.space/p/ainews-deepseek-v41-flash-763b-p8b</guid><description>We agree with Sebastian: this should have been DeepSeek v5</description><author>Latent</author><pubDate>Sat, 12 Sep 2026 05:56:05 +0000</pubDate></item><item><title>Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking</title><link>https://arxiv.org/abs/2609.10657</link><guid isPermaLink="true">https://arxiv.org/abs/2609.10657</guid><description>arXiv:2609.10657v1 Announce Type: new Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of…</description><author>Arxiv</author><pubDate>Sat, 12 Sep 2026 04:00:00 +0000</pubDate></item><item><title>An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics</title><link>https://arxiv.org/abs/2609.10712</link><guid isPermaLink="true">https://arxiv.org/abs/2609.10712</guid><description>arXiv:2609.10712v1 Announce Type: new Abstract: We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised…</description><author>Arxiv</author><pubDate>Sat, 12 Sep 2026 04:00:00 +0000</pubDate></item><item><title>Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills</title><link>https://www.marktechpost.com/2026/09/11/anthropic-adds-plugin-evals-to-claude-code-6-grader-types-a-no-plugin-baseline-and-a-ci-gate-for-skills/</link><guid isPermaLink="true">https://www.marktechpost.com/2026/09/11/anthropic-adds-plugin-evals-to-claude-code-6-grader-types-a-no-plugin-baseline-and-a-ci-gate-for-skills/</guid><description>Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded. It answers 3 questions plugin developers…</description><author>MarkTechPost</author><pubDate>Fri, 11 Sep 2026 21:05:55 +0000</pubDate></item><item><title>Y Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too</title><link>https://techcrunch.com/2026/09/11/y-combinators-garry-tan-wants-u-s-open-weight-ai-labs-to-distill-frontier-models-too/</link><guid isPermaLink="true">https://techcrunch.com/2026/09/11/y-combinators-garry-tan-wants-u-s-open-weight-ai-labs-to-distill-frontier-models-too/</guid><description>Tan argues that frontier models themselves trained on public human knowledge so access to capable AI should be "a form of public good."</description><author>Techcrunch</author><pubDate>Fri, 11 Sep 2026 20:59:47 +0000</pubDate></item><item><title>Kimi-maker Moonshot AI targets $2B in annual revenue</title><link>https://techcrunch.com/2026/09/11/kimi-maker-moonshot-ai-targets-2-billion-in-annual-revenue/</link><guid isPermaLink="true">https://techcrunch.com/2026/09/11/kimi-maker-moonshot-ai-targets-2-billion-in-annual-revenue/</guid><description>While K3's usage figures have declined slightly in recent months, OpenRouter data currently shows as many as 300 billion tokens being generated each day by K3 models on the system.</description><author>Techcrunch</author><pubDate>Fri, 11 Sep 2026 19:35:54 +0000</pubDate></item><item><title>Fugu Max and Fugu Ultra v2 — Sakana's router splits into cheap and strong</title><link>https://sakana.ai/fugu-max-release/</link><guid isPermaLink="true">https://sakana.ai/fugu-max-release/</guid><description>Sakana AI released Fugu Max and Fugu Ultra v2, two versions of its orchestrator that routes each task to other models. Fugu Max costs $2 and $6 per million tokens; Fugu Ultra v2 scores 48.3 on Chartography against Opus 5's 27.3.</description><author>Sakana AI</author><pubDate>Fri, 11 Sep 2026 12:00:00 +0000</pubDate></item><item><title>SenseNova-U1.5 report — the recipe behind SenseTime's 8B unified model</title><link>https://arxiv.org/abs/2609.11929</link><guid isPermaLink="true">https://arxiv.org/abs/2609.11929</guid><description>SenseNova-U1.5 is SenseTime's 8B-MoT unified multimodal model, and its technical report is now public. The paper covers an encoder-free, VAE-free design that reads and generates images at native resolutions up to 4K.</description><author>SenseTime</author><pubDate>Thu, 10 Sep 2026 12:00:00 +0000</pubDate></item><item><title>GPT-Live-1 in the API — full-duplex voice for $0.05 a minute</title><link>https://developers.openai.com/api/docs/models/gpt-live-1</link><guid isPermaLink="true">https://developers.openai.com/api/docs/models/gpt-live-1</guid><description>GPT-Live-1 is now in the OpenAI API. The full-duplex voice model listens and speaks at once, handles interruptions, and hands deeper reasoning to the models and tools you pair it with. Voice sessions cost $0.05 per minute.</description><author>OpenAI</author><pubDate>Thu, 10 Sep 2026 12:00:00 +0000</pubDate></item><item><title>North Small Translate — Cohere's open translation model beats DeepL on WMT26</title><link>https://cohere.com/blog/north-small-translate</link><guid isPermaLink="true">https://cohere.com/blog/north-small-translate</guid><description>North Small Translate is a 218B-parameter open-weights translation model from Cohere Labs with 25B active parameters. It scores 83.60 on WMT26 across all languages, ahead of DeepL NextGen at 81.37.</description><author>Cohere</author><pubDate>Thu, 10 Sep 2026 12:00:00 +0000</pubDate></item><item><title>SWE-2 — Cognition's coding model lands within a point of Fable 5.1</title><link>https://cognition.com/blog/swe-2</link><guid isPermaLink="true">https://cognition.com/blog/swe-2</guid><description>SWE-2 is Cognition's new coding model, post-trained from Moonshot's 2.8-trillion-parameter Kimi K3. It scores 50.0% on FrontierCode 1.1 Main against 50.9% for Fable 5.1, and Cognition says it costs 64% less to run at that score.</description><author>Cognition</author><pubDate>Thu, 10 Sep 2026 12:00:00 +0000</pubDate></item><item><title>DeepSeek V4.1 Flash — a 552B open-weight rebuild with 1M context</title><link>https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash</link><guid isPermaLink="true">https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash</guid><description>DeepSeek V4.1 Flash is now generally available under the API name deepseek-flash, with MIT open weights on Hugging Face. It scores 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, ahead of Opus 5.0 and GPT-5.6 Sol.</description><author>DeepSeek</author><pubDate>Thu, 10 Sep 2026 12:00:00 +0000</pubDate></item><item><title>YuE2-3B — open music model tops Suno v5 and v6 on WildSongBench</title><link>https://huggingface.co/m-a-p/YuE2-3B</link><guid isPermaLink="true">https://huggingface.co/m-a-p/YuE2-3B</guid><description>YuE2-3B turns lyrics and a style prompt into a full 48 kHz song with vocals. It writes an editable melody-and-chord score first, then renders the audio, and scores 6.96 on WildSongBench against Suno v6's 6.56.</description><author>Multimodal Art Projection</author><pubDate>Wed, 09 Sep 2026 12:00:00 +0000</pubDate></item><item><title>NCP-ArchPreview — an 8.9B model that predicts concepts, not just tokens</title><link>https://arxiv.org/abs/2609.10715</link><guid isPermaLink="true">https://arxiv.org/abs/2609.10715</guid><description>NCP-ArchPreview is an 8.9B open-weight language model from Shanghai AI Lab that predicts multi-token "concepts" alongside normal next-token prediction. It reaches OLMo-3-7B's final pretraining loss using 51.3% of the training tokens.</description><author>Shanghai AI Lab</author><pubDate>Wed, 09 Sep 2026 12:00:00 +0000</pubDate></item><item><title>Show-Harness — one semantic interface lets a VLM drive a robot arm</title><link>https://arxiv.org/abs/2609.10522</link><guid isPermaLink="true">https://arxiv.org/abs/2609.10522</guid><description>Show-Harness is an open control layer that exposes a robot as discrete semantic action units a vision-language model can reason over. Show Lab at NUS shipped the Apache-2.0 code, six LoRA adapters and the demonstration data with the paper.</description><author>Show Lab, NUS</author><pubDate>Wed, 09 Sep 2026 12:00:00 +0000</pubDate></item><item><title>Suno v6 — a music model family trained only on licensed catalogues</title><link>https://suno.com/blog/introducing-v6</link><guid isPermaLink="true">https://suno.com/blog/introducing-v6</guid><description>Suno v6 is a new family of music models built with data licensed from Warner Music Group, BMG and Believe. It ships as v6, v6-wild and v6-mini, and Suno says it will retire its older models as the rollout finishes.</description><author>Suno</author><pubDate>Wed, 09 Sep 2026 12:00:00 +0000</pubDate></item><item><title>Gander — an open 9B model that listens, watches and works at once</title><link>https://github.com/Omni-Interaction-Gander/Omni-Interaction-Agent</link><guid isPermaLink="true">https://github.com/Omni-Interaction-Gander/Omni-Interaction-Agent</guid><description>Gander is an open 9B omni-interaction model that takes streaming video, speech and text together. You can interrupt it mid-sentence, and a separate reasoning agent keeps working on long tasks in the background.</description><author>Tencent Hunyuan Speech Team</author><pubDate>Wed, 09 Sep 2026 12:00:00 +0000</pubDate></item><item><title>Nex-N2.5 — three open-weight agent models, up to 1.6 trillion parameters</title><link>https://github.com/nex-agi/Nex-N2.5</link><guid isPermaLink="true">https://github.com/nex-agi/Nex-N2.5</guid><description>Nex-N2.5 is a family of three open-weight agent models from Nex AGI. The 1.6-trillion-parameter Max tier scores 92.6 on BrowseComp and 86.1 on Terminal-Bench 2.1. All three are Apache-2.0, and mini and Pro run free on OpenRouter.</description><author>Nex AGI</author><pubDate>Tue, 08 Sep 2026 12:00:00 +0000</pubDate></item><item><title>AuK — Tencent's open speech model generates and edits audio by instruction</title><link>https://github.com/Tencent-Hunyuan/AuK</link><guid isPermaLink="true">https://github.com/Tencent-Hunyuan/AuK</guid><description>AuK is an MIT-licensed 1.5B speech model from Tencent Hunyuan, Shanghai Jiao Tong University and the Shanghai Innovation Institute. It generates and edits speech from plain-language instructions. Code and weights shipped on September 8, 2026.</description><author>Tencent Hunyuan</author><pubDate>Tue, 08 Sep 2026 12:00:00 +0000</pubDate></item><item><title>NeoHorse-1 — open 4B and 9B models post-trained by a routing harness</title><link>https://arxiv.org/abs/2609.08183</link><guid isPermaLink="true">https://arxiv.org/abs/2609.08183</guid><description>NeoHorse-1 is a pair of Apache-2.0 models, 4B and 9B, that TokenRhythm post-trained by routing agent work across a pool of models and turning the results into training data. The 4B scores 64.87 on the team's ten-benchmark average, up from 58.94.</description><author>TokenRhythm</author><pubDate>Tue, 08 Sep 2026 12:00:00 +0000</pubDate></item></channel></rss>