Skip to main content

How Does Storm Process Audio With Multiple Speakers?

Learn how Storm handles group recordings, toolbox talks, and meetings where multiple people are speaking, and how to get the best results.

Written by Nina Yang

Storm's AI Form Fill captures audio as a single stream from one device. When multiple people are speaking — in a meeting, toolbox talk, briefing, or site walkthrough — Storm processes everything it hears and attempts to fill the form based on the full conversation. Understanding how this works helps you get better, more accurate results.


How Storm Captures and Processes Audio

Storm records a single audio stream from the device microphone. It does not receive separate audio channels per participant, and it does not have access to a meeting attendee list. This is different from dedicated meeting tools like Fathom, Otter, or Fireflies, which connect inside the meeting platform and receive a separate audio feed for each participant.

Tool type

Audio input

Speaker identification

In-meeting tools (Fathom, Otter, Fireflies)

Separate audio feed per participant + attendee list

High confidence — knows who said what

Storm (and similar on-device tools)

Single audio stream from one device mic

Speaker labels generated for multi-speaker recordings; accuracy improves when names are used verbally

Because Storm works from a single audio stream, it generates speaker labels by distinguishing voices by audio patterns rather than separate channels. Accuracy improves when speakers identify themselves or others by name during the recording.


What Storm Can and Cannot Do

Storm can:

  • Ingest and transcribe audio with multiple speakers

  • Process group discussions, toolbox talks, briefings, and meeting minutes

  • Infer context and decisions from the full conversation

  • Generate speaker labels when multiple speakers are detected

Storm cannot (yet):

  • Create or assign Actions to other people — any Actions created via AI Form Fill are attributed to the person doing the recording


Best Practices for Multi-Speaker Recordings

1. Name People Explicitly

This is the single most important thing you can do. Storm infers speaker identity from context, so using names makes attribution far more accurate.

Instead of

Try

"You install the safety barriers."

"Nina to install the safety barriers."

"She'll follow up on that."

"Sarah will follow up on the permit by Friday."

"He's responsible for the induction."

"James is responsible for running the site induction."

Pronouns like you, he, she, and they give Storm no way to determine who is being referred to. Using a person's name directly resolves this.

2. Avoid Crosstalk Where Possible

When multiple people speak at the same time, audio quality drops and Storm's ability to transcribe accurately is reduced. Where practical:

  • Have one person lead the recording (e.g. the meeting chair or site supervisor)

  • Ask others to speak one at a time, especially when capturing decisions or actions

3. Recap Key Decisions and Actions Out Loud

A brief spoken summary at the end of the recording helps Storm capture what matters most — even if parts of the conversation were unclear. For example:

"To recap: Nina to install safety barriers by Thursday. James to complete the risk assessment by end of week. No further actions."

4. Fill in the Gaps After Processing

For any speaker attribution or action items that Storm gets wrong or leaves blank, review and edit the form after processing. You can re-run AI Form Fill with additional notes or corrections if needed.


How Multi-Speaker Audio Works With Other Form Inputs

Storm processes audio alongside all other inputs available in the form at the time of recording. In a group discussion or toolbox talk, this means:

  • Weather field — if weather has already been captured in the form before the recording, Storm reads it as context automatically. No one needs to mention conditions out loud.

  • Existing form content — dates, timestamps, previously entered notes, dropdown selections, and table content already in the form are all read by Storm and used to anchor what it hears in the recording.

  • Instructions vs. content — in group discussions, someone may direct Storm mid-recording (e.g. "summarise this as meeting minutes" or "add another row to the workers table"). Storm attempts to distinguish these instructions from actual form content, but being explicit and having one person lead the instructions helps accuracy.


Current Limitations

  • Storm generates speaker labels by distinguishing voices by audio patterns from a single audio stream. Accuracy can be affected by crosstalk, similar voice tones, or background noise — dense or overlapping discussions may still require manual review.

  • Using names verbally during the recording still significantly improves attribution accuracy.

  • Actions created via AI Form Fill are attributed to the person doing the recording, not the person assigned in the meeting


FAQ

Can Storm tell who said what in a group recording?
Storm generates speaker labels when multiple speakers are detected, distinguishing voices by audio patterns in the transcript. Because Storm works from a single audio stream without a participant list, it cannot confirm speaker names from voice alone. Using names explicitly during the recording still significantly improves attribution accuracy.

Is this a Storm limitation or an AI limitation generally?
It applies to any AI tool that captures a single audio stream from one device — including tools like Granola. It is not unique to Storm. Dedicated meeting tools (Fathom, Otter, Fireflies) avoid this by connecting inside the meeting platform where they receive separate per-participant audio feeds.

Can I assign Actions to other people using AI Form Fill?
Not yet. Action creation via AI Form Fill is coming soon. When available, any Actions created will be attributed to the person doing the recording. To assign actions to others, name them explicitly in the recording so they're captured in the form content, then create and assign the Action manually.

What's the best setup for recording a toolbox talk or briefing?
Have one person hold the device and lead the recording. Speak names when assigning tasks. Do a short verbal summary of decisions and actions at the end before stopping the recording.

Did this answer your question?