Get Started
Sound Classification
On this page
  1. Commands
  2. Settings
  3. Output Variables
  4. JSON Output

Sound Classification

A type of the audio analyzer. Classifies the sound from the microphone input (or the speaker output) into one of the 521 AudioSet categories learned by Google's YAMNet model (speech, music, dog barking, alarms, glass breaking, screaming, car horns, and more). Used for abnormal sound detection and environmental sound analysis. The input is chosen in the analyzer setting Audio Source: Microphone (Input) or Speaker (Output).

Note Audio analysis writes results only after the Start Analysis command is run following Add Analysis.

Commands

CommandDescription
Specific Sound?Checks whether the target sound is detected at or above the threshold.
Detection ResultWrites the name and probability of the most probable sound.
All ResultsWrites every recognition result, ordered by probability, as JSON.

Settings

SettingDescription
Target SoundThe sound watched by Specific Sound?. One of the 521 categories (English names); the default is Speech.
Detection Threshold (%)Probability at or above this value counts as the sound. 10 to 100, default 50.

Output Variables

FieldTypeValue
Sound NameTextEnglish name of the most probable sound.
Sound ProbabilityNumberProbability of the recognized sound, 0 to 100.
Sound DetectedDigitalTrue when Specific Sound? finds the target sound at or above the threshold.
JSON ResultTextArray of every recognition result.

JSON Output

An array with, per result, label and confidence (0 to 100), ordered from the highest probability. Unlike Action Recognition and Hand Motion Recognition there is no index key.

[
  { "label": "Speech", "confidence": 85 },
  { "label": "Music", "confidence": 10 }
]

For the full list of the 521 categories see YAMNet's official class map (yamnet_class_map.csv).