Home / News / Gemini 3.8 Live & Musk Safety Push

News

Gemini 3.8 Live & Musk Safety Push

US

Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

Google DeepMind released its most advanced live dialogue and speech-to-speech models on September 15. Gemini 3.8 Live prioritizes scale, cost efficiency, fluid dialogue, near-real-time visual grounding, mid-conversation language switching across 97 languages, and background tool/API execution. The Extended Thinking variant emphasizes multi-step reasoning and high-complexity tasks, topping the Artificial Analysis Speech-to-Speech Quality Index at 82.6, leading agentic benchmarks such as τ-Voice (68.6%) and τ-Voice-banking (35.1%), and scoring 97.7% on Big Bench Audio. Audio is SynthID-watermarked. Models are available via Gemini API, AI Studio, Search Live, Gemini Live, and Workspace tools for subscribers, with pricing at $0.005/min input and $0.018/min output.

OpenAI holds early talks on funding round valuing the company at roughly $1.2 trillion or higher ahead of IPO

Investors initiated discussions for a private round that could lift OpenAI’s valuation from its March $852 billion post-money figure (after a $122 billion raise) to about $1.2 trillion, with some reports citing internal views closer to $1.5 trillion based on Codex traction and models including GPT-6 Astra. Annualized revenue recently exceeded $40 billion. CEO Sam Altman has indicated no 2026 IPO due to safety work. Talks remain preliminary.

Agility Robotics unveils Digit 5 humanoid designed for barrier-free work alongside humans

The 5'11", 284 lb Digit 5 lifts 50 lb (40% more than prior), reaches 7.2 ft, runs 90 minutes on a charge that takes 9 minutes (enabling >20 hours work per day), and uses AI/sensors for cooperative safety: autonomously avoiding, stopping, or squatting when detecting people. Swappable end-effectors support varied tasks. Early access targeted for first half of 2027, general availability by end of 2027; company holds >$300 million in multi-year orders and plans European expansion. Built on 65,000+ hours of real-world Digit 4 data.

Gates Foundation pledges at least $1 billion over two years to expand global AI access

Funding targets education (40%, including tutoring), healthcare (40%, diagnostics and decision support), agriculture (10%, smallholder advice), and digital foundations (10%, non-English datasets). The Goalkeepers report warns market-driven AI risks widening inequality unless deliberately directed toward underserved populations. Bill Gates noted governments lag in preparedness.

Salesforce launches Koa CRM reasoning model on NVIDIA Nemotron and expands Claude integration

At Dreamforce, Salesforce announced Koa, its first reasoning model post-trained on Nemotron 3 Super using synthetic data modeled on 27 years of CRM workflows (no customer data). It matches or exceeds leading models on CRM Bench with 3x fewer errors. Available in pilots now, general availability expected winter 2026 in US regions. Separately, Salesforce in Claude (beta on paid plans) brings 37 sales skills and live CRM actions under existing permissions via MCP; AIforce enables broader interfaces including Slack.

OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind

OpenAI’s global affairs chief stated the company has engaged Anthropic and DeepMind on safety for weeks, aligning with earlier calls for standards. Concurrently, Trump administration figures dismissed existential risk concerns as overblown while emphasizing competition with China.

Meta’s Mark Zuckerberg states labs have sufficient incentives via competition and liability to build safely

In an X post and statements, Zuckerberg argued individual labs should pace development as needed, citing Meta’s earlier delay of Muse agent for security, and favored independent evaluators over coordinated slowdowns.

Elon Musk proposes cross-lab “test harness” for safety evaluations

At the All-In Summit, Musk urged xAI/SpaceXAI, OpenAI, Anthropic, Google, Meta, and leading Chinese firms to allow rivals to run safety tests on each other’s models before release, checking for capabilities such as bioweapon assistance or deception, rather than self-grading.

Non-US West

EU Commission President Ursula von der Leyen to invite frontier AI labs for talks on pacing risks

In her annual address, von der Leyen announced plans to convene leading labs on supporting industry efforts to manage frontier risks, citing the EU AI Act’s positioning of Europe to shape global standards. The move follows researcher resignations and safety warnings.

China

China releases AI Safety Governance Framework 3.0 and calls for true multilateralism under UN auspices

Foreign Ministry spokesperson Guo Jiakun highlighted the updated framework, which revises risk classifications and recommends technical and comprehensive safeguards. Beijing emphasizes equal focus on development and security, human control, and opposition to “overstretching” national security concepts that privilege one country’s interests. People’s Daily commentary rejected AI as a “monopoly of great powers” and criticized dual standards on model distillation.

China’s Ministry of State Security chief warns AI poses risks to Party rule, infrastructure, and military competition

Chen Yixin detailed threats including foreign models enabling espionage, cyberattacks, and misinformation, while calling for tighter Party oversight and accelerated domestic security systems. Official statements frame US slowdown calls as containment tactics.

Non-China East

Limited highly notable, major-impact developments meeting the threshold emerged in the window from independent East Asian sources outside China; ongoing chip and infrastructure activity continues without discrete breaking announcements of comparable global consequence in the period.

Grok / xAI / Elon Companies (AI, Robotics, Supporting Tech)

Elon Musk posts on Grok Bot Galaxy livestream and Grok Build upgrades

Musk amplified the three-day Grok Bot Galaxy event (September 15–17) in which SpaceXAI staff use Grok Bot agents to build a company from a blank slate in real time, covering engineering, sales, support, and marketing workflows. Separate posts highlighted Grok Build v1.0.33 improvements to MCP structured JSON, memory management, long-running session reliability, subagent cancellation, and edge-case fixes for heavy agent workflows. Musk also endorsed connecting specialized MCP tools to Grok Bot for deeper research engagement and suggested Grok for problem-solving contexts.

Musk urges leading labs including Chinese firms to peer-test models via test harness

At All-In Summit, Musk proposed that xAI/SpaceXAI, OpenAI, Anthropic, Google, Meta, and top Chinese companies grant rivals pre-release access for independent safety evaluations focused on high-risk capabilities.

No discrete new Optimus or FSD capability announcements of major impact appeared in the window; focus remained on agent automation demonstrations and cross-lab safety proposals.

Summary

US advancements in live multimodal models, specialized enterprise reasoning, safer humanoid deployment, philanthropic access expansion, and massive capital inflows demonstrate continued leadership in translating frontier capabilities into practical, scalable systems that can accelerate productivity and global problem-solving. Shortfalls in coordinated safety governance and regulatory clarity risk allowing competitive pressures to outpace robust evaluation frameworks, potentially amplifying misuse vectors if peer-review mechanisms remain voluntary and incomplete.

European moves to convene labs and leverage existing regulatory frameworks reflect a measured institutional approach to risk management that prioritizes structured dialogue without immediate disruption to development pipelines. Shortfalls in rapid deployment of concrete evaluation infrastructure or binding timelines leave the region more reactive than proactive relative to the pace of capability gains elsewhere.

Chinese framework updates and multilateral rhetoric project an image of responsible governance while accelerating domestic capabilities and open-weight efforts. These steps, however, appear calibrated primarily to protect political control and close the technology gap rather than independently constrain capability growth, raising concerns that safety measures may serve competitive positioning more than universal risk reduction. Shortfalls in transparent, independently verifiable evaluation processes and the explicit prioritization of catching up leave open questions about whether internal controls can reliably contain advanced systems once deployed at scale.

Steady technical progress in regional infrastructure and specialized models continues without dramatic frontier leaps or governance shifts that would redefine global trajectories. Any shortfalls in independent safety leadership or rapid scaling of novel architectures remain secondary to broader supply-chain and talent dynamics.

Continued emphasis on practical agent autonomy for business processes and calls for competitive safety testing underscore efforts to operationalize advanced models while addressing coordination gaps. Progress in agent reliability tools positions these systems for broader automation roles, though the absence of immediate binding peer-review agreements leaves verification dependent on voluntary participation.

Frontier model releases, specialized enterprise and robotics systems, large-scale capital and philanthropic commitments, and intensifying safety discourse across major players illustrate accelerating capability deployment alongside unresolved governance coordination challenges. The period highlights practical utility gains in voice agents, CRM reasoning, humanoid collaboration, and agent-driven workflows, set against divergent national approaches to risk management that continue to prioritize competitive positioning.

Term

Domain

The name you type into a browser, like yourname.com. You rent it from a registrar. It is the address people put on an invoice, not the website files themselves.

Acronym

DNS

Domain Name System

The internet's address book. When someone types your domain, DNS tells their computer which server actually holds your site or your mail.

Acronym

MX

Mail Exchange

A DNS record that says which service should receive email for your domain. If MX is wrong, mail to you@yourdomain bounces or lands in the wrong inbox.

Acronym

IMAP

Internet Message Access Protocol

The usual way a phone or laptop mail app talks to your mailbox. Mail stays on the server, so the same inbox shows up on every device.

Term

Registrar

The company that rents you the domain name and lets you point it somewhere. Cloudflare Registrar and Porkbun are registrars. Squarespace or Google can also act as one if you bought the name through them.

Term

Nameserver

The computers, named by the registrar, that answer DNS questions about your domain. Changing nameservers is how you hand DNS from one host to another.

Acronym

2FA

Two-factor authentication

A second check after the password, usually a code from an authenticator app or a hardware key. SMS codes are weaker. You type 2FA. An agent does not.

Term

Stack

The small set of tools you actually run the company on: mail, invoicing, calendar, a site. On this site, a stack page is one job with a named pick and a named skip.

Term

Agent

Here, an AI helper that can draft files, checklists, and clicks you can undo. It does not hold passwords, 2FA, the card, or the final words on a public page.

Acronym

CRM

Customer relationship management

Software for tracking people you sell to: notes, follow-ups, a pipeline. A spreadsheet or starred mail is enough until those follow-ups stop fitting in a table.

All news