Moonshot shops a 30% cut. DeepSeek books $70 million.
Kimi K3 on Azure, AWS, and GCP. Plus India’s 9,000 Vera Rubin order.

China’s Moonshot AI is in early talks with Microsoft, Amazon, and Google to host Kimi K3, its 2.8-trillion-parameter open-weight model, on Azure, AWS, and Google Cloud, and is seeking up to 30 percent of the K3-related cloud-service revenue, Liam Mo and Fanny Potkin at Reuters report. Juro Osawa and Qianer Liu at The Information write that DeepSeek booked about 475 million yuan, or $70.7 million, of revenue in the first seven months of 2026 — roughly ten times its full-year 2025 take — even as a second funding round is underway. Saritha Rai at Bloomberg reports that Greenko-backed AM Intelligence, a Hyderabad AI-infrastructure firm, placed a binding order for 9,000 Nvidia Vera Rubin systems under an $8 billion plan to stand up a gigawatt of compute.
1. Moonshot shops a 30% cut for Kimi K3
Liam Mo and Fanny Potkin at Reuters report that China’s Moonshot AI is negotiating revenue-sharing agreements with Microsoft, Amazon, and Alphabet’s Google that would let those clouds host Kimi K3, the 2.8-trillion-parameter open-weight model — weights anyone can download — that few customers are likely to self-host because of compute cost. IPO-bound Moonshot, founded in 2023 by Carnegie Mellon-trained Yang Zhilin and backed by Alibaba, is seeking up to a 30 percent share of revenue generated from K3-related services on Azure, AWS, and Google Cloud, three people familiar with the talks said, in line with terms sources have said the startup has outlined for major customers.
Discussions are at an early stage and there is no certainty they will result in agreements, with the revenue split, data access, and how to audit token usage still unresolved. Moonshot did not respond; Microsoft, Google, and AWS declined to comment. Arena.ai ranked Kimi K3 first on a web-interface-building benchmark, Artificial Analysis says performance is comparable to OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.8 on complex multi-step tasks, and the company raised more than $2 billion in May as it prepares for a potential Hong Kong listing, sources have said — even after U.S. officials last month floated a blacklist and accused it of stealing from Anthropic’s Fable, a charge Moonshot has denied.
2. DeepSeek books $70.7 million, a tenfold jump
Juro Osawa and Qianer Liu at The Information report that DeepSeek, the Hangzhou AI lab, generated about 475 million yuan, or $70.7 million, in revenue in the first seven months of 2026, roughly tenfold its full-year 2025 revenue, according to two people with knowledge of the company’s financial data. The same people said DeepSeek recorded a net loss of about 715 million yuan, or about $106 million, from January through July, compared with 935 million yuan, or about $139 million, for all of 2025.
The Information’s public teaser says a second funding round is underway and that the revenue growth could give a boost to those talks and to a potential initial public offering down the line, without naming a date, a round size, or a valuation.
3. India orders 9,000 Vera Rubin systems
Saritha Rai at Bloomberg reports that AM Intelligence, a Hyderabad AI-infrastructure company backed by the group that owns renewable-energy producer Greenko Group, has placed a binding order for 9,000 Nvidia Vera Rubin systems, Nvidia’s next rack-scale computing platform, under a broader $8 billion plan to stand up a gigawatt of computing capacity. Servers equipped with the systems are slated to come online next year in southern India, the company said, for customers that include major cloud-service providers, AI labs, and organizations aiming to develop homegrown Indian models.
Mahesh Kolli, founder and president of Greenko Group, told Bloomberg the initial capacity has already been purchased by a U.S. customer he declined to identify because of a nondisclosure agreement, and that the group will offer capacity in India, the United States, Finland, and Malaysia, with undersea cables letting U.S. companies use the compute at about 300 milliseconds of latency. He said the group aims to reach 5 gigawatts of data-center capacity by 2030 across India, Europe, and elsewhere.
4. Wrtn raises $72.2 million at more than $722 million
Joyce Lee at Reuters reports that Wrtn Technologies, the South Korean AI-services platform, has raised around 100 billion won, or $72.2 million, in a Series C that valued the company at more than 1 trillion won, or $722 million. The round brings total fundraising to about 230 billion won, or $166 million, and adds Coreline Ventures and Eugene Asset Management alongside existing backers including Goodwater Capital and Korea Development Bank.
Wrtn said OOC, its North America-focused AI entertainment platform, surpassed 10 billion won, or $7.22 million, in monthly revenue within three months of its May launch, and it expects total 2026 revenue to exceed 200 billion won, compared with 47.1 billion won last year. The company said it will use the funding to expand overseas and to tighten safeguards for teenage users of Crack, its AI entertainment platform, including time-use limits, stronger parental-consent procedures, and lower spending caps.
5. Gates says there is no plan
Bill Gates, in a GatesNotes memo, writes that the AI era will be among the most turbulent times in human history and that we are not preparing adequately. Jeffrey Dastin at Reuters reports that Gates has policy ideas he wants to discuss with China’s Xi Jinping. Lindsay Ellis at The Wall Street Journal calls the 5,784-word warning “There Is No Plan,” and Cristina Criddle at the Financial Times writes that Gates is calling for “human reserved” jobs to protect the labor force.
Karen Weise at The New York Times writes that tech executives are privately “very worried” about AI disruption but publicly downplay the risks to protect fundraising and planned IPOs. In an MIT Technology Review interview with Mat Honan, Gates says the industry has already crossed thresholds on bio-capabilities, cyber-capabilities, psychosocial capabilities, job-market destruction, and even lack of control, puts bioterrorism at about 50 times more likely than a natural pandemic, and floats a token tax, a robot tax, and human-reserved jobs — adding that the industry will not self-regulate because “we’re trying to raise trillions.”
6. Z.AI confirms Ox Alpha, weights tonight
Luz Ding at Bloomberg reports that China’s Z.AI Co., also known as Zhipu, confirmed that Ox Alpha — the new model that swept to the top of online usage charts with high performance at zero cost — is a new iteration of its GLM series and said it will release the weights tonight. Techmeme adds that Ox Alpha topped OpenRouter’s leaderboard, the public ranking of which models developers actually send traffic to.
7. Mechanical Turk shuts down September 30
Annie Palmer at CNBC reports that Amazon Web Services will close Mechanical Turk, the 21-year-old crowdsourced work platform Jeff Bezos once called “artificial artificial intelligence,” effective September 30, 2026, following an assessment. Launched in 2005 to farm out tasks that were easy for humans and hard for computers — labeling data, transcribing audio, answering surveys — the marketplace at one point served more than 500,000 workers.
Krista Pawloski, an organizer at Turkopticon, the worker-advocacy group that tracks Mechanical Turk conditions, told CNBC the platform had been in decline as Amazon seemed to invest fewer resources and as Scale AI, Mercor, and Prolific entered. A 2023 Swiss study cited in the same story found that up to 46 percent of Mechanical Turk workers used AI models to complete the tasks they were paid to do by hand.
Watch
Dwarkesh Patel — who will control the world’s usable FLOPs
Dwarkesh Patel is joined by Dylan Patel, founder of SemiAnalysis, the research shop that models chip and data-center economics, for a deep dive into lab economics and who will control most of the world’s usable FLOPs, the floating-point operations that actually train and run models. Dylan’s argument is that OpenAI and Anthropic started the year around two gigawatts of power each and will end it above five, so about a third of this year’s new compute already terminates at those two labs, and the contracts already signed point to 40 to 50 percent of next year’s incremental watts.
The reason they can outbid everyone is the revenue-per-megawatt gap: serving GPT-4 on Hopper lost money, while serving GPT-5.6 or Opus 5 and Fable 5 now clears well past the $10 million to $15 million-per-megawatt base cost, with Anthropic as high as $50 million, which they plow back into training. Dylan’s non-consensus call is that the labs will keep moving compute from inference toward research as that internal return beats selling tokens, because revenue adds have plateaued while new megawatts have not, and the rest of the hour is the crowding-out — $11 trillion of AI capital expenditure modeled through 2029, maybe $5 trillion of it credit.
The one move is to pick one workflow you currently run on a frontier API and write down what happens to it if inference capacity at the labs gets scarcer and more expensive over the next 18 months, then decide on paper whether you need a second path — open weights, a router, or owned GPUs — before the bid-up shows up in your invoice.
NLW — the AI model tier list
NLW, host of The AI Daily Brief, walks through a weekend model tier list — a ranking of which AI models belong in your stack — and why routing, sending each job to the cheapest model that can still do the work, now beats picking a single lab. The list that started the argument put Anthropic’s Fable 5 alone in S-tier and OpenAI’s GPT-5.6 Sol in A, and the useful part is the paradox: Fable is the genius nobody wants to work with and nobody wants to fire, while Sol is the slightly dumber robot that does exactly what you tell it, which is still the default you would miss more.
The serious numbers sit under the ranking. AT&T already serves 40 percent of employee queries with open models and is targeting 60 to 70 percent while holding OpenAI and Anthropic spend flat, a router cut coding cost 56 percent with a 2 percent quality drop, and Vercel’s gateway flipped from 28 percent open-weight tokens to 62 percent in two months. Gavin Baker’s end-state, which NLW flags, is closed labs keeping 60 to 90 percent of the economic value on 15 to 25 percent of the tokens.
The one move is to run the same job this week on your default frontier model and on a cheaper or open alternative, write the quality delta and the bill, and only then decide whether a router or a second provider belongs in the stack.
Basil Chatha — what actually works in production voice AI
Basil Chatha hosts the pilot of Latent Space’s Forward Deployed series with five people who ship voice agents — Basia Sudol, head of enterprise solutions at Decagon; Varun Singh, chief product and technology officer at Daily; Steven Diaz, forward-deployed-engineer manager at Vapi; Tyler D'Silva, founding forward-deployed engineer at Retell AI; and Sudarshan Kamath, founder of Smallest AI — for a deep dive into what actually works in production voice AI in 2026. The consensus from the room is not the demo: nobody serious is shipping real-time speech-to-speech in production yet, and the cascaded pipeline of speech-to-text, a language model, and text-to-speech is still what you can guardrail, debug, and tool-call against.
The hard problems they actually fight are latency versus intelligence, waterfalls of fallback models for when a provider dies mid-call, turn-taking that can tell a think-pause from the end of a sentence, models that read the first and last four percent of a giant prompt and forget the middle, and unit economics that now get compared directly to human labor. Outbound is easier than inbound, getting an executive to like the voice can be harder than any model choice, and HIPAA and CRM tool calls without dropping the customer are the last mile.
The one move is to write the pipeline on one page as speech-to-text, then the model, then text-to-speech, with a named fallback model, a latency budget in the neighborhood of 800 milliseconds round-trip, and a hard cap on prompt size, and not to switch to speech-to-speech until you can observe a failed tool call.
Peter Yang — building better AI evals with Claude Code
Peter Yang is joined by Hamel Husain and Shreya Shankar, the instructors behind an evaluations course that has now taught more than 4,500 students, for a deep dive into building better AI evals with Claude Code — tests that score a model’s outputs the way a product owner would. They split the craft in two: top-down criteria are what a domain expert would demand in a vacuum, length and actionability and structure, which models are decent at inventing, and bottom-up criteria are the failure modes you only find by reading real outputs, which Shreya says Claude is very bad at inventing.
Her error-discovery skill in Claude Code spends about 15 minutes building a custom review interface over your traces, lets you annotate open-ended notes in place while a monitor tool clusters those notes into a rubric, then fans the rubric out as pass-or-fail judges. Hamel’s companion finding is that Braintrust Loop, Arize, LangSmith, and coding agents all recover the obvious tool-call failures about equally well and systematically miss the taste-and-judgment failures that actually differentiate a product, and auto-eval precision in their test sat around 80 to 90 percent, which means a tenth to a fifth of the flagged “errors” are red herrings if you accept them blindly.
The one move is to dump 10 to 20 raw inputs plus the outputs you actually shipped into a folder, invoke the error-discovery skill, label until the failure modes saturate, then promote those modes into yes-or-no judges instead of one giant vibes prompt.
Max Junestrand — learning the legal market faster than anyone else
At Y Combinator Startup School 2026, Max Junestrand, co-founder and CEO of Legora, the agentic operating system for lawyers now used across 50 countries, sits with YC partner Gustav Almströmer for a deep dive into learning the legal market faster than anyone else — the path from a May 2023 rejection to $1 million and then $100 million in annual recurring revenue in about 18 months. The founding team were not lawyers: they cold-emailed attorneys off firm websites and offered to pay the hourly rate for lunch, and Max’s line is that you need the willingness to learn the market, not a law degree on the cap table. They bet against the 2022 idea that you had to fine-tune a proprietary legal model and instead routed general models, later splitting traffic so expensive frontier models handle hard litigation, where tokens are cheap next to partner time, and cheaper open-source models handle consumption-heavy work.
The muscle he wants other founders to copy is evaluations. Legora hired lawyers whose job mixed customer work with building use cases for those tests, kept an internal “Legora Bench” until a wave of new models made publishing it interesting, and treats the ability to evaluate new models and new use cases as core intellectual property for routing. After a $35 million Benchmark and Redpoint round there was a month when interest income exceeded customer revenue, so they froze sales for six months because lawyers give you one chance in a demo.
The one move is to write a one-page evaluation for the workflow you sell — three tasks where you would pay frontier rates and three where open weights are enough — and refuse to demo the product until that split is real, not a slide.
