[AUDIO AI DAILY NEWS RUNDOWN] Grok 4.6 Storms Frontier, Autonomous AI Hacks Taiwan, and Meta’s Data-for-Discounts Trade (August 13, 2026)
🎧 Listen ADS-FREE: https://podcasts.apple.com/us/channel/djamgamind/id6760446113
Visit our Research Hub at https://djamgamind.com/pdfs
Summary: In today’s briefing, we analyze “Autonomous Cyber Offensives, The Enterprise Data Bargain, and Physical Intelligence Portability.” We deconstruct SpaceXAI’s release of Grok 4.6 and its 24/7 Grok Bot workspace. We examine the world’s first fully autonomous AI cyberattack on Taiwanese government infrastructure, alongside President Trump’s memo authorizing private sector counter-hacking. We break down Meta’s Muse Code contributor pricing model, IBM’s massive enterprise partnership with OpenAI, Google DeepMind’s Gemini Robotics 2 full-body humanoid model, and China forcing Meta to unwind its $2 billion acquisition of Manus.
Important Topics:
-
SpaceXAI Launches Grok 4.6 & Grok Bot: SpaceXAI releases Grok 4.6, scoring 61 on the Intelligence Index at $2/$6 per 1M tokens. Elon Musk previews Grok 4.7 for release in 3 to 4 weeks. SpaceXAI also debuts Grok Bot, placing autonomous agent teams on dedicated virtual machines for $200/month.
-
Autonomous AI Agents Attack Taiwan Infrastructure: State-linked hackers deploy eight autonomous agents (using Hermes and OpenClaw) to map 21 Taiwanese government systems and compromise 2,500 personnel records, marking the first known fully autonomous government cyberattack.
-
Trump Authorizes Private Sector “Hack Back” Policy: A White House memo allows vetted US private security firms (posting a $1M bond) to conduct offensive cyber operations against foreign cybercrime networks targeting US victims.
-
Meta Launches Muse Code with Data Contributor Discount: Meta introduces Muse Code and Muse Spark 1.2, offering a “contributor tier” ($0.10 input / $0.20 output per 1M tokens) that gives Meta training rights over developer prompts, outputs, and repository workflows.
-
IBM Partners with OpenAI for Enterprise Modernization: IBM creates a dedicated OpenAI Practice with thousands of certified consultants to integrate GPT-5.6, Codex, and Project Daybreak cyber defenses into legacy enterprise architectures.
-
Google DeepMind Unveils Gemini Robotics 2 (GR2): GR2 controls full-body humanoid movement (legs, torso, arms, and 22-degree-of-freedom hands) across multiple robot chassis using a single set of trained weights.
-
China Forces Meta to Unwind $2B Manus Deal: Chinese regulators (NDRC) order Meta to divest from AI agent startup Manus due to cross-border foreign investment and export control restrictions.
-
Twitch Introduces AI Training Opt-Out Toggle: Twitch adds a “Training for Generative AI” setting allowing creators to exclude streams, clips, and chats from Amazon’s model training sets.
-
Gamgee Launches YC-Backed Canine Cancer mRNA Vaccines: Founder Paul Conyngham launches Gamgee, using ChatGPT and AlphaFold to design personalized mRNA cancer vaccines for dogs following successful trials on his dog Rosie.
-
Google Gemini App Crosses 1 Billion Users: Google’s Gemini app hits 1 billion monthly active users, with 63% of interactions occurring via voice commands.
-
MiniMax Releases H3 HD Video Model: MiniMax drops H3, a 33B parameter multimodal video model setting new records on video editing leaderboards, though requiring US/EU/UK/SK users to apply for custom licenses.
🔗 RESOURCES
-
AI Learning App Recommendation: AI & ML Tutor PRO https://apps.apple.com/ca/app/ai-ml-tutor-pro/id1610947211
-
DJAMGATECH: Carrer Booster - Master AWS, Azure, AI & GCP Certifications | https://apps.apple.com/ca/app/djamgatech-ai-cert-exams-prep/id1560083470
-
DJAMGAMIND KIDS Bedtime Adventures:
⚗️ PRODUCTION NOTE: We Practice What We Preach.
AI Unraveled is produced using a hybrid “Human-in-the-Loop” workflow.
SpaceXAI’s Grok 4.6 storms the frontier
Image source: SpaceXAI
The Rundown: SpaceXAI just launched Grok 4.6, its new top model with benchmark scores that compete with frontier options like Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol while charging 60% less — a massive turnaround for the once-ridiculed model.
The details:
-
Grok 4.6 scored a 61 on Artificial Analysis’ Intelligence Index, passing Sol and now sitting behind just Opus 5 (63) and Fable 5 (62).
-
The release edges Fable on real-world professional and legal work, while beating Sol on a pair of coding benchmarks that 4.5 had trailed.
-
4.6 also undercuts the frontier on price at $2/$6 per 1M tokens, with Elon Musk calling it “objectively No. 1 when considering intelligence, speed & cost”.
-
Musk also revealed that Grok 4.7 “should be ready in 3 to 4 weeks”, saying the upgrade will “exceed all current models”.
Why it matters: Is Elon about to overtake the frontier? The Cursor/xAI partnership had already shown results, but Grok has now gone from punchline to the frontier. With 4.7 on the way, this is the type of pressure that might force the hands of frontier rivals with both new models and price — something not many would’ve expected a year ago.
Grok takes on OpenAI in coding LINK
-
xAI has released Grok 4.6, its latest AI model, which the company says matches OpenAI’s GPT-5.6 Sol on several coding and knowledge-work benchmarks, marking Grok’s return to the front of the AI race.
-
xAI credits its partnership with and possible purchase of coding company Cursor, whose real-world usage data trained Grok 4.5 and, after a longer training run, Grok 4.6, which launches first inside Cursor alongside the Grok Build coding agent.
-
Grok still faces reputation problems, from its role in making non-consensual nude images to weak business uptake, with just 4% of companies using AI tools paying for xAI, according to Ramp’s AI Index.
Trump lets private firms hack back LINK
-
President Trump signed a White House memo letting vetted private security firms apply for government approval to hack foreign cybercrime groups behind ransomware, phishing, fraud, and scam operations that targeted U.S. victims.
-
The National Coordination Center will run the program, with executive directors from the Justice and Homeland Security departments overseeing operations and each participating company posting at least $1 million in bond or escrow.
-
Firms must stop immediately and alert the center if they detect activity beyond approved limits, such as accidentally targeting U.S. citizens or systems; Americans reported losing over $20.8 billion to cyber crime in 2025.
Twitch trains Amazon AI on your streams LINK
-
Twitch now lets users turn off a setting that allows their content to help train Amazon’s generative AI models, meaning streams, VODs, clips, chats, and channel pictures and text stay out of future training.
-
The “Training for Generative AI” toggle sits under the Security and Privacy tab, and it was switched on by default when The Verge checked, though Amazon hasn’t confirmed whether that’s the standard setting.
-
Opting out doesn’t stop other AI-supported tools like captions, AutoMod safety, and recommendations, and if you chat on someone else’s stream, that person’s opt-out choice decides whether the chat gets used.
AI agents nearly autonomously hacked Taiwan LINK
-
Hackers believed to be linked to China ran what researchers call the first fully autonomous cyberattack on a government, using free AI tools to break into Taiwanese systems, steal over 2,500 personnel records, and compromise 85 accounts.
-
The four-day campaign in early July deployed up to eight AI agents at once, built on two open-source frameworks called Hermes and OpenClaw, which mapped 21 government systems before targeting Taiwan’s nuclear safety agency and energy firms.
-
According to Israeli firm Dream, the tool devised its own attacks, ranking and switching paths when techniques failed, and bypassed the model’s safeguards by posing the intrusion as an authorized penetration test rather than a real attack.
Apple may pay publishers for Siri news LINK
-
Apple is talking with publishers about paying them to let its updated Siri assistant pull current news and information, according to a Wall Street Journal report on the company’s plans.
-
Instead of the usual fixed licensing fee tied to broad content access, Apple has pitched a pay-as-you-go model that compensates publishers each time their material is actually used by Siri.
-
Apple has weighed a nine-figure budget for these payments, and the deals come as the company works to improve Siri, which is expected to launch later this year.
Flock tightens rules after police misuse LINK
-
Flock Safety, the police camera company that logs where Americans drive, said Thursday it will add new safeguards to curb officer misuse, including changes to how long it keeps people’s vehicle location data.
-
By default, Flock will now hold movement data for one week instead of one month, and officers wanting to keep records longer must enter a case number to justify the extended access.
-
The company is fully rolling out Audit Assistance, a tool tested earlier in 2026 that locks out users whose searches show “abnormal activity,” though the ACLU called the reforms minor and still wants a 48-hour retention limit.
Dog that got an AI cancer vaccine now has a company
Image source: Gamgee
The Rundown: Paul Conyngham introduced Gamgee, a Y Combinator-backed startup building mRNA cancer shots matched to the mutations in a single dog’s tumor, coming just months after he made one with ChatGPT and AlphaFold for his own dog Rosie.
The details:
-
Several vet visits over 11 months missed Rosie’s mast cell cancer, and once surgery, chemo, and immunotherapy all failed, her outlook fell to just months.
-
Conyngham used AI tools for the science and analysis, with UNSW then developing a personalized vaccine that helped reduce Rosie’s tumors.
-
Sam Altman met with Conyngham in March following the initial story, saying it had him thinking that “this should be a company”.
-
Australian vets are hosting the first trials, with Gamgee taking cases worldwide and managing the custom design, treatment, and monitoring end-to-end.
Why it matters: Saving dogs is a mission we can get behind, and Conyngham’s story is an amazing example of being able to just do things in the AI age. The hope is also that the start of personalized cancer treatments doesn’t end at dogs, and AI is going to be a major driver of pushing the same sort of research forward into the human realm.
IBM just gave OpenAI a shortcut into the enterprise
As forward deployed engineers become the hottest new role in AI, OpenAI is trying forward deployed partners.
On Thursday, the AI lab announced a partnership with IBM to embed its technology into IBM’s consulting arm. Through the partnership, IBM’s forward-deployed units, trained through the OpenAI Partner Network, will help enterprise clients modernize their workflows with OpenAI’s tech.
Additionally, IBM will join OpenAI’s “elite partner tier,” and will launch a dedicated OpenAI Practice that includes thousands of consultants and engineers, specifically trained with OpenAI certifications.
Colleen Kapase, OpenAI’s vice president of global partnerships and ecosystem, exclusively told The Deep View that partnering with IBM represents a “significant” opportunity. “IBM works with enterprises around the world and across virtually every major industry, including many of the complex and highly regulated environments where AI can have enormous impact.”
The partnership has three areas of focus:
-
Updating legacy operations with OpenAI’s frontier models, including GPT-5.6 and products like Codex and ChatGPT Work, across domains including finance, procurement, customer operations, and HR
-
Modernizing legacy applications and software development using Codex and ChatGPT Work, aiming to help enterprises simplify engineering processes and speed up product shipment
-
And increasing IBM’s participation in Project Daybreak by combining OpenAI’s cyber capabilities with IBM Autonomous Security, its agentic security platform, to help clients manage both cyber and AI model risk
The goal of the partnership is to help IBM Consulting enterprise clients get beyond one of the major hurdles standing in the way of successful AI deployments: The fact that so much enterprise work relies on fragmented legacy systems and tools that aren’t well fitted for the modern and fast-paced nature of AI, Andy Baldwin, global senior vice president of IBM Consulting, told The Deep View.
“Most organizations don’t struggle with getting access to AI tools or models anymore. The real challenge is putting AI to work securely at scale and effectively across large, complex organizations,” said Baldwin. “At the same time, AI is changing the threat landscape. Organizations are facing more sophisticated, AI-enabled threats and growing pressure to strengthen and build agentic first cyber defenses.”
Muse Code Wants Your Data
Meta will cut coding bills from dollars to pennies for developers who let the company learn from their work.
What’s new: Meta introduced Muse Code, a command-line agentic coding harness, and Muse Spark 1.2, the capable, low cost-per-task model behind it.
-
Input/output: Text, images, video, and PDF in (up to 1,048,576 tokens), text out
-
Features: Adjustable reasoning levels (none, minimal, low, medium, high, xhigh), tool use, structured output, web search, context caching, background subagents that persist across a session
-
Performance: Achieved 57 points on Artificial Analysis’ Intelligence Index, first on Vals AI’s Finance Agent v2 and on Artificial Analysis’ AA-LCR
-
Availability/price: Muse Code in beta for macOS and Linux, Muse Spark 1.2 via Meta Model API. Standard tier $1.25/$0.15/$4.25 per million input/cached/output tokens (prompts and outputs not used for training); contributor tier $0.10/$0.002/$0.20 per million input/cached/output tokens (prompts and outputs used for training); web search $2.50 per thousand queries
-
Undisclosed: Parameter count, architecture, knowledge cutoff, training data, and method details
How it works: Muse Code runs in a terminal. Given a software task, it plans changes, writes code, and checks results at each step using Muse Spark 1.2. Meta trained the model to work with Muse Code, using data from Muse Spark 1.1. The company describes three design choices that distinguish the agent from a single loop that calls a model repeatedly.
-
A main agent delegates to a set of subagents that remain for the length of a session instead of being created and discarded for each task. Subagents edit in parallel inside isolated worktrees, separate working copies of a repository that keep simultaneous changes from conflicting, Mark Zuckerberg wrote.
-
Because they persist, subagents retain context they have already learned about a repository rather than re-derive it. The subagents also determine when to report back to the main agent on their own.
-
The agent writes every model call, tool run, plan approval, and file edit to a log on the user’s machine. If the agent crashes, it reads the log and resumes from the step it reached rather than starting over, which helps agents to work on long-running tasks.
-
The agent ships with three default skills that users can call: /plan converts a request into a roadmap the user must approve, /grill probes that plan for weak points, and /goal drives work until an objective is met.
Performance: Independent evaluations place Muse Spark 1.2 a notch below the intelligence frontier, but at a lower cost per task than most models around or above its level.
-
On Artificial Analysis’ Intelligence Index, a composite of nine evaluations of economically useful tasks, Muse Spark 1.2 set to xhigh reasoning (57, $0.40 per task) ranked sixth, above Grok 4.5 set to high reasoning (56, $0.36 per task), and just behind Qwen3.8-Max set to reasoning (58, $1.13 per task). The new model’s Intelligence Index score was 4 points above last month’s Muse Spark 1.1 (53, $0.29 per task).
-
On Artificial Analysis’ AA-LCR, a test of reasoning across long documents, Muse Spark 1.2 set to xhigh reasoning (83.3 percent) outperformed all other models tested.
-
On the Vals Index, a composite of finance and coding tasks weighted by potential economic impact, Muse Spark 1.2 set to xhigh reasoning (71.88 percent, $0.70 per task) ranked fifth, ahead of Claude Opus 4.8 set to max reasoning (70.36 percent, $7.52 per task) and behind GPT-5.6 Sol set to max reasoning (73.12 percent, $7.46 per task) — less than one tenth of the cost per task.
-
On Vals AI’s Finance Agent v2, which assigns models the work of entry-level financial analysts, Muse Spark 1.2 set to xhigh reasoning (60.60 percent, $0.77 per task) ranked first of 45 models, significantly cheaper than second-place Claude Opus 5 set to max reasoning (58.63 percent, $5.12 per task) and taking roughly half the time per test.
Behind the news: Muse Spark 1.2 is already an inexpensive model, but if it catches on — OpenAI and other companies have tried similar initiatives — a contributor discount for the model’s use in Muse Code is potentially a transformative one. Companies’ appetite for training data drives new policy pushes and business intiatives.
-
In a long essay, Mark Zuckerberg argued that United States labs are disadvantaged by restrictions on training data.
-
Of all training data, high-quality code is a scarce commodity. Hugging Face published The Stack v3 last week, a 4.9 trillion-token crawl of public GitHub assembled to replace its years-old predecessor. But a public repository shows only polished code, not all the reasoning, mistakes, and repairs that went into it. Sessions from a working coding harness capture that entire process.
-
The Muse Spark 1.2 discount lands in a price competition that has escalated over the summer: OpenAI cut GPT-5.6 Luna’s prices by 80 percent to $0.20/$1.20 per 1 million tokens of input/output, and DeepSeek-V4-Flash-0731 arrived at $0.14/$0.28 per 1 million tokens of input/output. Meta’s contributor tier undercuts them both and virtually everyone else selling high-performing models through an API.
Why it matters: The contributor tier buys Meta something its apps don’t supply. Facebook, Instagram, and WhatsApp generate enormous quantities of data, but not the kind of coding data required to train coding agents. Meta is short of such data and is willing to give away most of the price of the Muse Spark 1.2 to get it. Meta sells the same model at two prices, and the discount buys Meta the right to train on whatever passes through the agent. That trade requires no contract: A developer picks it by typing a different model name. The tier containing these data terms caps at 100 requests per minute per team versus 3,000 requests per minute for the standard tier, limits that make it practical mainly for individuals and small teams — the developers least likely to have a lawyer on retainer, and the ones whose entire product may sit in the repository the agent reads.
We’re thinking: No one is forced to give up their data to use Meta’s best model or agent. But developers weighing Muse Code’s discount are deciding, whether they think about it deliberately or not, what their own code and expertise are worth. Model builders have long trained on developers’ code by scraping it from public repositories and forums. Meta is trying to turn that knowledge transfer into a market. Like all markets, this one rewards the side that knows what its goods are worth, and the goods for sale here, both repositories (public or private) and a recording of how the work was done, lacked a clear price before. Meta has transparently priced the discount. Likewise, developers should price the value of their data before calling the trade a true bargain.
Google’s Robotics Model Has Legs
Google’s latest vision-language-action model can walk a humanoid robot across a room, crouch to a low shelf, and close a five-fingered hand around a light bulb. Google says it’s the first model in its robotics family to run the legs and the hands from one set of weights, a breakthrough that simplifies model training and design.
What’s new: Google released Gemini Robotics 2 (GR2), a model that turns camera images and typed instructions into commands for a robot’s joints. Whereas earlier models in the Gemini Robotics family drove only the upper body for tabletop work, GR2 moves a humanoid’s legs, torso, arms, and hands together. One fixed set of the model’s trained weights runs three machine setups across two robot bodies.
-
Input/output: Camera images and text instructions (input) to joint commands (output), driving either a five-fingered hand with 22 degrees of freedom (Sharpa or Inspire hands) or a simple two-fingered gripper (Robotiq)
-
Features: Leg and hand control; multi-embodiment (one checkpoint runs Apptronik’s Apollo 2 humanoid with Sharpa hands, the same Apollo 2 with Inspire hands, and a Franka Duo arm rig with a Robotiq gripper)
-
Performance: Self-reported success rates of 45.7 to 76.3 percent on whole-body pick-ups, 32 to 92 percent on multi-finger tasks, and 74.2 to 89.6 percent on gripper tasks
-
Availability: Early-access partners, with no public API
-
Undisclosed: Google published no model card for Gemini Robotics 2, and gave no base model, parameter count, training-data breakdown, or price
How it works: The release pairs Gemini Robotics 2 with Gemini Robotics ER 2, a separate reasoning model that breaks a job into steps and hands them to GR2 one at a time. Google published no model card or technical paper for GR2, so its safety report and announcements provide what we know.
-
Google trained the model for each task in the demonstration videos using a mix of teleoperation, in which a person operates the robot remotely while the movements are recorded, plus video examples and simulation.
-
GR2 runs one set of trained weights across three machine setups; no version was trained per configuration. Motion transfer, introduced alongside Gemini Robotics 1.5, trains a single model on data pooled from robots of different shapes, sensors, and joint counts. This makes data recorded on one machine setup valuable learning material for data on another.
Performance: Google ran every evaluation of GR2 on its own defined tasks and hardware, and no outside group has published tests of the model. Google reports the whole-body and gripper figures as averages over several tasks in a category and the finger figures as individual tasks. The nearest baseline is Google’s own first robotics model from March 2025, which was adapted to a two-armed Franka robot and averaged 63 percent across the tasks it was given.
-
An Apollo 2 robot with Inspire hands, trained to pick things up while using the legs, succeeded 76.3 percent of the time from a shelf, 68.4 percent from a table, and 45.7 percent from the floor. Google reports each as an average over several tasks in that category.
-
Fine finger work using Apollo 2 with Sharpa hands produced the lowest figures Google published: Tying a trash bag succeeded 44 percent of the time, sealing a zippered bag 40 percent, and sweeping with a dustpan 32 percent.
-
On the Franka robot, GR2 averaged 89.6 percent on precise insertion, 78.9 percent on tool kitting, and 74.2 percent on general pick and place. This task list differs slightly from that measured for Gemini Robotics 1.
-
Gemini Robotics ER 2’s judgment about whether a task was physically possible for the robot depended on being told what Gemini Robotics 2 had practiced. With no summary of that training it was right 62.0 percent of the time; with a highly detailed summary, 95.8 percent.
-
Google also introduced a safety benchmark, ASIMOV-Agentic. Google found that no system it tested, its own included, can both catch hazards to a person from a moving robot and avoid stopping for nothing. Holding needless stops under 5 percent meant missing more than 40 percent of the moments a person was too close. Consequently, Google recommends running these models alongside conventional physical safety equipment rather than in place of it.
Screwing in a light bulb: One pair of tasks in Google’s results illustrates how an apparently simple task, when reversed, can be a challenge in robotics. Unscrewing a bulb succeeded 92 percent of the time, the highest figure in Google’s finger-work set. Screwing one in succeeded 36 percent of the time. Unscrewing starts from a settled position (bulb already in the socket), so the hand needs only grip and rotation. Screwing one in has to establish bulb-in-socket alignment first, with the hand wrapped around the bulb.
Behind the news: An updated Gemini Robotics reasoning model called ER 2 is paired with this release. ER 2 plans steps, tracks progress from a video feed, and calls an action model as a tool. Another model, Gemini Robotics On-Device 2, runs on the robot’s own hardware without a network connection and adapts to an unfamiliar two-armed body in a few hours using typically fewer than 200 examples.
Why it matters: Many robotics models are trained for one task, one machine, and one environment at a time, and changing a single variable often means starting from scratch. There is also no corpus of robot motion at anything close to the scale of the text and images behind vision-language models. The latest Gemini Robotics models are an experiment in learning from more generalized training data than has been available for robots in the past. It shouldn’t matter if you change the type of hands, number of joints, or tasks performed: The goal of a true multi-embodied robotics model is to successfully transfer knowledge from one setup to another, as transformer-based language models have done so successfully with text.
We’re thinking: How many Gemini robots does it take to screw in a light bulb? At the current 36 percent success rate, just under three.
MiniMax’s State-of-the-Art Video Model Is Only Minimally Open
A free-to-download model sets a new standard for video generation and editing, but its license comes with unexpected restrictions.
What’s new: MiniMax released H3, a high-definition video generation model that accepts a wide range of input media. But the model’s license requires users in the U.S., the UK, the European Union, and South Korea to submit an application to MiniMax to use it under the same terms as the rest of the world. Key components of the model also remain proprietary, at least for now.
-
Input/output: Up to twelve files including text (7000 characters), images (30 MB per file), audio (15 MB), and video (50 MB) in, video (up to 2000 pixels wide and 15 seconds long) with audio and text captions out
-
Architecture: Transformer, 33 billion parameters; separate text, audio, and video encoders, preprocessor, 2k video upscaler
-
Features: Video and audio editing, multi-shot output, six aspect ratios
-
Performance: First on Artificial Analysis’ video editing leaderboard, second (or tied for first, within the margin of error) in text-to-video and image-to-video
-
Availability/price: Weights free for noncommercial and commercial uses under MiniMax H3 license; via MiniMax’s API at $0.13 per second for 2K resolution, $0.08 per second for 768p resolution output, input costs vary from free for audio to $0.04 per image and $0.13 per second for high-definition video; prompt re-generation module costs $0.90 per million tokens of input and $3.60 per million tokens of output.
-
Undisclosed: Exact parameter count, training data, technical report
How it works: H3’s architecture consists of three modules — a contextual processing system, a video/audio generation base model, and a high-definition upscaler. Only the base model is free to download, and it comes with restrictions.
-
MiniMax says it trained H3 on “real, natural data” to ensure data quality and scalability. Unlike MiniMax’s earlier video models, all audio types (voice, music, and sound effects) were trained together and are modeled through a single encoder. The team trained the model’s reference and editing functions using natural language rather than fixed presets. The team sought to combine different media types (sound, video, etc.) earlier in the process.
-
Supported generation modes include text-to-video, first- or last-frame image to video, and image, audio, and video references, plus any combination. For example, a user prompt might instruct the model to begin with a single still image, reference the camera movement in two videos, and the audio score from a third, with text instructions for how the scene should unfold.
-
Input is first processed by H3-Context-IR. This reasoning module interprets the prompts and input media and generates a new text prompt instructing the base model how to blend them together. Users of the base model alone need to be explicit in their instructions for each media type, or use their own preprocessing system; users of the full API pipeline benefit from the Context-IR module doing most of that work.
-
The base module generates 768p video, that is, 1792 pixels by 768 pixels (given a maximum 21:9 aspect ratio). Three encoders process text, video, and audio, respectively; a separate variational autoencoder also encodes video.
-
A third module, H3-Regenerate-2K, instructs the base module to regenerate the 768p video in native 2K resolution (2000 pixels wide and up to 4667 pixels long). The optimized prompt also includes the original media references and user prompt as context, allowing the regenerated video to add missing details rather than only extrapolating from the source video.
-
The weights’ license requires commercial users to prominently display the model name on any product using H3 and bars all users from distilling another model on H3’s output. It also prohibits use that may harm minors, interfere with elections, or violate local law. The license also identifies the United States, United Kingdom, European Union, and Republic of Korea as “excluded territories” and requires users of the weights in these territories to apply for a license “to ensure [their] usage is lawful, responsible, and without infringing any rights.”
Behind the news: MiniMax H3’s head-to-head human-preference ELO scores put it squarely in a top three with Google’s Gemini Omni Flash and Bytedance’s Dreamina Seedance 2.0. Black Forest Labs’ FLUX models have typically included an open weights release for developers, but its recent FLUX 3 model is proprietary and API-only. It’s yet to be independently tested, but Black Forest Labs’s tests suggest it would be a fourth model vying to be state-of-the-art. Dreamina Seedance 2.5 also awaits testing.
Why it matters: Despite the restrictive license and unusual territorial restrictions, MiniMax H3 is clearly a top video generation model, offering commercial users a strong and versatile set of tools for video generation on par with Google’s much more expensive competitor. We also get a glimpse of how a top video model works under the hood. It appears that MiniMax’s strategy of foregoing synthetic data — especially difficult when good video transcripts are hard to come by — and bundling prompt optimization into the pipeline is paying off.
We’re thinking: Even in this paranoid period of AI development, it makes no sense to restrict use of weights by country beyond uses that would break that country or region’s laws. It violates the fundamental meaning of openness: the idea that anyone can download a model, see how it works, and put it to use. We hope these restrictions don’t become a trend.
AI Can Help Heal Romantic Distress
Chatbots that are designed to treat mental health issues typically require multiple sessions, posing a risk that users will drop out before they receive much benefit. Researchers showed that chatbots can provide relief in a single session.
What’s new: Thomas Menzel at Technical University of Munich and University of Cambridge, along with Michel Schimpf and Thomas Bohné at University of Cambridge, built overit, a chatbot app designed to help users recover from romantic breakups. In a randomized, controlled trial, users felt substantially better after one conversation.
Key insight: Romantic breakups can continue to cause distress long after they occurred. Often, the former lovers are troubled by self-limiting beliefs about themselves (“I was abandoned because I am not enough”) or the world (“nobody could want me”). According to memory reconsolidation theory, recalling a distressing memory, contemplating a self-limiting belief about it, and then presenting an interpretation that contradicts the earlier belief (“You did the best you could, but you were failed by someone you trusted”) can durably update the painful memory. If that’s true, a chatbot that elicits a self-limiting belief and guides the user toward a counterfactual interpretation could bring about a lasting reduction in distress.
How it works: Users filled out a survey that included a breakup distress score, breakup timing, and the ex-partner’s name (along with follow-up surveys after the trial period). Then they discussed their breakups with a mobile app based on Claude Sonnet 4.5.
-
Given their survey responses, the app guided them through four phases of conversation. The app (i) asked open-ended questions about a breakup and its impact, (ii) elicited beliefs and identified at least one self-limiting belief, (iii) offered alternative perspectives, and (iv) asked users what they had learned from the conversation and how they felt now. It progressed from one phase to the next after a number of turns or reaching a certain milestone, capping conversations at 18 turns.
-
With each user input, the app asked Claude Sonnet 4.5 to assess the current phase by checking the last three turns against five milestones: (i) identifying a self-limiting belief, (ii) challenging it, (iii) steering the user toward a counterfactual interpretation, (iv) articulating a new insight, and (v) concluding the dialog.
-
Then the model considered the conversation history, instructions for the current phase (for example, to identify the core self-limiting belief in phase two), and survey data and generated a response.
Results: The authors ran a randomized, controlled trial with 171 participants in the U.S. and UK who had experienced a breakup within 18 months on average. Half of participants conversed with the app, via text or voice, in a single conversation of roughly 20 minutes; the other half did not interact with the app. The authors measured the participants’ distress via the Breakup Distress Scale, a 16-item questionnaire that yields scores between 16 and 64.
-
After 7 days, the group that used the app had experienced a large reduction in distress (from 35.3 to 26.6) relative to the control group (from 35.9 to 32.2).
-
After one month, the group that used the app showed lower distress (26.0) than the control group (29.0).
-
People who used the app were more likely to report a “sudden insight” about their breakup (such as “I had been carrying the blame for my failed relationship instead of recognizing that I did everything the best I could and I was failed by someone I trusted”): 61.7 percent versus 19.3 percent. Those who had such experience tended to report feeling better afterward.
Why it matters: Chatbots frequently sycophantically affirm a user’s expressions. This can be troublesome in therapeutic settings, where an analytical approach would be more helpful. This work addresses that issue by dividing each input into calls for evaluation and generation. Prompting the LLM first to evaluate the conversation’s state and then to generate a response separates the tasks of tracking therapeutic progress and expressing empathy to the user, which supports the LLM’s ability to provide helpful output. This approach offers a template for building goal-directed conversational agents that challenge users (in this case, to examine and reinterpret painful memories) rather than simply comforting them.
We’re thinking: Any chatbot session is less expensive than a human therapist, but a single 20-minute conversation that can meaningfully improve a subject’s mood and attitude is truly favorable economics, even at Claude’s rates.
Anthropic adds AI watermarks to Claude outputs
Image source: Images 2.0 / The Rundown
The Rundown: Anthropic just published a support page detailing plans to embed invisible watermarks in text, code, and file outputs from Claude, putting the EU AI Act’s transparency rules into practice globally across its user base.
The details:
-
Text and code generated by new Claude models will carry a hidden watermark that carries through when copied and pasted outside the platform.
-
Files will get the same C2PA provenance standard label that is already used across the industry to mark AI-generated media.
-
Older Claude models will need to be ‘retrofitted’ to include the labelling, with models shipping after Aug. 2 having the tool built in.
-
Anthropic is also planning to release its own detection tools and noted that a mark shows content was “processed by Claude”, not necessarily fully authored.
Why it matters: The move is a polarizing one, with takes ranging from positive to dystopian and overreaching. Anthropic isn’t alone in signing the AI Act, with OpenAI, Google, Meta, and others (notably absent: xAI) likely needing to navigate the issue as well. For the negative camp, the case for private and open models continues to grow.
Grok Bot is the latest AI agent teammate build
Image source: xAI
The Rundown: SpaceXAI and Cursor just rolled out Grok Bot, a new beta app that turns Grok into an iMessage-style chat of teammates that have their own dedicated computers, work autonomously, and can access everyday apps just like a real user.
The details:
-
Bots are able to coordinate with each other independently or in group chats, and learn from workflows and improve over time.
-
New specialist agents can be generated by Bots inside a job or receive handoffs from another worker, with work continuing 24/7 without a laptop open.
-
Bots can also access websites and apps on a user’s behalf and observe and create workflows to automate tasks.
-
The beta covers iPhone and Mac, Windows, and Linux desktops for SuperGrok Heavy and top Cursor tiers, with Elon Musk saying access will expand soon.
Why it matters: The group chat is becoming a top form factor for running teams of agents. OpenClaw and Hermes already live in messaging apps, and Jack Dorsey’s Buzz built a Slack rival where agents are channel members. Grok Bot feels like a middle-ground version built for the masses (if the price comes down significantly).
xAI co-founder’s River banks $1.1B for personal AI
Image source: River AI / The Rundown
The Rundown: Former xAI co-founder Igor Babuschkin just raised $1.1B for River AI, his two-month-old startup focused on building open-source AI that is trained, tuned, and controlled by the individual instead of a major corporation.
The details:
-
Before xAI, which he left a year ago after helping co-found the startup with Elon Musk in 2023, Babuschkin worked at OpenAI, Tesla, and Google.
-
River claims its live API, released at launch, turns open-weight models into ones a business actually can own, cutting training to just minutes.
-
Babuschkin told the NYT the goal is a highly customizable assistant that follows across devices while running on the user’s own private hardware.
-
He also said, “We don’t want these AI companies to rule the world and control this superpowerful technology,” hitting at closed U.S. frontier labs.
Why it matters: It’s another eye-popping raise for a company without much to show yet, but Babuschkin’s pedigree is worth betting on. With questions about access (Fable/Mythos ban) and forced watermarks (above) putting privacy and control top of mind in the industry, River is pulling at a string with a whole lot of interest.
Gemini hits 1 billion monthly users LINK
-
Google’s Gemini app has passed 1 billion monthly users, a milestone reached less than a month after the company reported 950 million users, putting it closer to OpenAI’s ChatGPT, which hit the same mark in May.
-
Google said 63% of people now speak to Gemini instead of typing, one in five interactions involve screen sharing or live camera feeds, and Apple users account for over 100 million of the app’s monthly active users.
-
The company said Gemini generates 150 million images each day, a tool growing popular with small businesses, and noted that macOS users prompt roughly twice as often as people on other platforms.
SpaceXAI launches Grok Bot LINK
-
SpaceXAI, the SpaceX division once called xAI, has launched an early beta of Grok Bot, agents that sign into a user’s apps and websites to finish real work rather than just answering prompts.
-
Each Bot runs on its own computer around the clock, learns routines by watching users demonstrate tasks, remembers preferences, and can hand jobs to other Bots, including a “Chief of Staff” Bot that directs specialist ones.
-
Available today for macOS, Windows, Linux and iOS, Grok Bot costs $200 a month for individuals through Cursor Ultra and $120 per seat monthly for teams, with organizations otherwise pointed to a waitlist.
China forces Meta to unwind Manus deal LINK
-
Manus said it will soon go back to running as an independent company after Chinese regulators ordered Meta to unwind its $2 billion purchase of the AI agent startup, ending a deal announced in December.
-
China’s National Development and Reform Commission issued the order in April, investigating whether the acquisition broke rules on foreign investment amid tighter export controls on cross-border tech deals during the US-China AI race.
-
As part of the split, Manus told some users they must back up any data they created on or after December 29, 2025, the date Meta first announced the takeover, to meet regional regulatory demands.
Zoom bug let callers hijack your device LINK
-
Researchers found flaws in Zoom that let attackers silently take over a target’s device during any screen-sharing call, hitting either the host or participants with no warning and no action needed from the victim.
-
The bugs affected Zoom on every operating system it supports, Windows, macOS, Linux, iOS, and Android, and the company put out a security advisory Tuesday with fixes it has already started rolling out.
-
A Security says it found the flaws in early June using freely available AI models, needing fewer than 20 prompts, work that its cofounder Omer Gull said would once have taken a five-person team roughly six months.
What Else Happened in AI on August 13th 2026?
Google announced that its Gemini app officially hit 1B users, saying the AI is now the fastest-growing product in the company’s history.
DeepSeek quietly rolled out V4 Pro, a new model coming in at just $0.435/$0.87 per 1M pricing with benchmarks that slot slightly below both the frontier and open leaders.
AI coding startup Lovable secured $400M in new funding that values the company at $13.3B, more than doubling its valuation from its previous December raise.
Google introduced the Pixel 11 lineup, embedding Gemini AI further across the devices, with new voice upgrades, multistep task capabilities, and proactive prompts.
OpenAI COO Brad Lightcap is leaving to “start something new,” saying there are “a few important new things the world will need to get right” as AI enters its next phase.
Anthropic reportedly told investors it plans to prioritize AI in healthcare and biology to help “mitigate some of the negative sentiment around AI”.
AI video company LTX released LTX-2.5, an open-weights world model that tops rivals like Gemini Omni Flash on internal tests for video generation speed and quality.
Nvidia released a small, fast Nemotron 3.5 Lightning model and anticipates its Nemotron 4 to be “on par with the best open-source AI models in the world.”