Image: Apple

Apple introduced Audio Intelligence on the Watch Series 12 and Ultra 4 at Wednesday’s hardware event. Sound Recognition can alert a wearer to sirens, alarms, doorbells, or a crying baby even when the iPhone is elsewhere. That feature is assistive. Live Rewind and Siri Recap are not. They convert nearby speech into text the wearer can keep. The person being transcribed did not press anything.

What the New Features Actually Do

Live Rewind holds a rolling 15-second audio buffer in a Secure Enclave on the watch. Double-press the Digital Crown and the last 15 seconds appear as a text snippet. The wearer can ask Siri about it or save it to a new Siri app. Apple says the raw audio never accumulates, is inaccessible to the operating system, apps, the user, and Apple, and is deleted if the paired iPhone is not available to complete the transfer.

Siri Recap uses ambient listening to generate high-level notes of conversations: a title, summary, and key points, not a verbatim transcript. It can be scheduled by time or location, left off, or switched from Control Center. Recaps auto-delete after seven days unless saved. Apple says speakers are not identified and that sensitive content such as financial details is filtered.

Activation of Live Rewind triggers an audible chime even on silent mode, plus a full-screen microphone animation. Both features are opt-in for the wearer. Live Rewind and Siri Recap arrive in English beta later this year and require a recent iPhone. They will not launch initially in the EU.

Holding up a phone to record is visible. Double-pressing a watch crown is not. A chime and a screen flash notify people nearby only if they are looking and listening. That is a notice to the room, not permission from each speaker. Device makers such as Plaud tell customers to obtain legally required consent before recording. Amazon’s Bee terms put compliance with minors’ privacy rules on the user, not the company. Apple’s design stores text, not audio. Courts have not settled how those snippets will be treated as evidence, or how two-party consent statutes apply when the original sound is discarded.

Apple’s privacy paper is more detailed than most wearable AI vendors offer. No stored recordings is a real design choice. It does not change the social fact: a conversation that used to die in the air can now exist as a searchable note on someone else’s wrist. Studies of recording devices and surveillance culture already show people change how they speak when they believe they might be captured. A subtler capture method widens that effect.

Why Apple Built It Anyway

Startups have spent two years selling pendants and pins that passively note a wearer’s day. Apple is answering that category from a product already on tens of millions of wrists. Sound Recognition is the cleanest use case. Live Rewind and Siri Recap are the competitive ones. The company can argue on-device processing and end-to-end encryption. It cannot argue that the other person in the conversation agreed to be summarized.

The watch is not “always recording” in the sense of an archive sitting on a server. It is always buffering speech long enough to freeze it as text. That distinction matters for Apple’s legal posture. It matters less to the colleague, parent, or stranger who just had 15 seconds turned into a note they never authorized. Normalization starts when the privacy-branded company ships the feature and calls the chime sufficient.