06
Proof followed agents into speech and the physical world.
Once an agent talks to a caller or controls a device, a plausible answer is a weak test. Teams need repeatable scenarios, device limits, recorded behavior, and visible approval points.
Physical operationsAnthropic previewed a shared hardware interface
Anthropic · August 27
The Model Hardware Standard is a model-agnostic research preview for programmable lab and manufacturing devices. Standardized drivers expose simple operations to agent harnesses through protocols including MCP.
Read the research preview →
Device safetyThe driver can carry the safety limit
Anthropic · August 27
MHS drivers can declare what a device measures, what can change, and which limits must be enforced. Anthropic is still building safety evaluations, approval practices, and a physical-safety roadmap with preview partners.
Review the safety design →
Voice evalsGoogle's simulator can speak to the agent
Google Developers · August 24
Google's ADK added live evaluation that drives a voice agent with a simulated user speaking audio. Teams can score multi-turn behavior against natural-language rubrics instead of trusting one clean demo.
Read the engineering guide →
Regression testingThe conversation can become a test artifact
Google Developers · August 24
The pipeline supports voices, language settings, turn caps, trajectory scoring, per-turn checks, CI execution, reconstructed transcripts, and playable audio. Synthetic users still need real-user coverage, especially for accents, interruptions, and unusual behavior.
See the evaluation workflow →
Operator read: Build the test around the consequence. For voice, test identity, interruption, tool use, and disclosure. For hardware, test state validation, hard limits, emergency stop, recovery, and the exact decision that requires a person.