Tech Note | Understanding Presenter Spotlight vs Speaker Spotlight

Learn the key differences between Presenter Spotlight and Speaker Spotlight to enhance your VisionSuite system effectively.

Updated at July 30th, 2026

Information


VisionSuite offers two distinct modes that serve fundamentally different production goals. Presenter Spotlight continuously tracks a single individual within a defined area using a dedicated vision-based tracking camera. Speaker Spotlight automatically frames whichever participant is currently speaking, using audio localization from a beamforming microphone and a rules engine to trigger camera presets. Choosing the right mode, or combining them depends on the room layout, production style, and the hardware available.

Presenter Spotlight

Presenter Spotlight assigns a single camera to the Presenter Tracking role. That camera runs the person-detection at 15 Hz, feeds detections into the global tracking system, and outputs a continuously smooth, responsive tracking shot. A deadband mechanism prevents the camera from reacting to minor movement. Adaptive zoom adjusts the field of view during initial shot acquisition so the presenter is professionally framed (full-body, American, half-body, or medium close-up).

 
 

Speaker Spotlight

Speaker Spotlight assigns multiple cameras to the Speaker Spotlight role. Instead of continuous vision tracking, the system reacts to voice-activity-detection (VAD) events from the room's beamforming microphone. When a participant begins speaking, VisionSuite provides a 3D source position with an uncertainty estimate. The audio shot handler converts that position into a framing target — computing the apparent person height from the speaker height in the configured shot container — then drives the camera to the calculated pan/tilt/zoom.

Because the camera image is not used for detection, the detections run at just 1 Hz. There is no deadband, no adaptive zoom, and no presenter-zone filtering. Camera movement is event-driven: the system moves only when a new active speaker is detected.

 
 

Comparison at a Glance

  Presenter Spotlight Speaker Spotlight
Primary Use Case Continuously follow one presenter on stage or at a lectern Automatically frame any active speaker across the room
Beamforming Mic Required No Yes
Trigger Vision - person detection in a zone Audio - VAD from microphone array
Detection Frequency 15 Hz (vision pipeline) 1 Hz (audio event-driven)
Zones

Not global - 2D zones drawn in each camera's FoV

 

Must have real camera hardware and VSA-100 to configure - cannot configure in emulation

 

Trigger, Exclusion, and Tracking

Global - 2D zones drawn on top of floorplan

 

Trigger Spotlight exclusion - does not exclude audio from microphones, only prevents active talker from triggering a rule

Deadband Yes - configurable size per shot container No - system can intelligently detect if an active talker is already framed to avoid unnecessary shot recalls and delayed camera switching
Adaptive Zoom Yes - during initial shot acquisition No - fixed shot calculation based on distance uncertainty and shot safety
Multi-person Handling

Tracks on VIP at a time

 

With a conductor camera, multiple people can enter zone to trigger static wide shot to capture more people

Switches to any new speaker automatically
Failure Mode VIP-lost triggers fallback static shot No speaker detected - overview on silence rule is triggered to switch to wide shot of room

Decision Guidance

The conditions below can be used to decide what implementation to use for your system.

Presenter Spotlight

  • The room has a defined stage, lectern, or presentation area where one person will be the primary focus.
  • Smooth, continuous camera motion is required - deadband produces broadcast-quality tracking.
  • The presenter often moves significantly (walking across a stage) and the camera must follow in real time.
  • You need zone-based constraints (e.g. ignore people shown on TV or behind glass wall).
 
 

Speaker Spotlight

  • The room is a meeting or collaboration space where any participant may speak.
  • There is no single "presenter" - the active speaker changes frequently and unpredictably.
  • You need to minimise system complexity - Speaker Spotlight requires fewer congifuration steps and no presenter zones.
 
 

Use Both

  • The room has both a defined presentation area and seated participants who need to be framed when they speak - for example, a lecture hall with audience Q&A, a boardroom with a lectern, or a hybrid training room.
  • You want the presenter to be tracked continuously on a dedicated camera while a second camera automatically frames any active speaker elsewhere in the room.
  • The production requires seamless switching between a presenter shot and audience/panel reaction shots without manual camera operation.
  • The production expects broadcast-style coverage where the presenter camera is always "warm" (tracking in the background) so the cut back is instant, when a room speaker stops talking.