Appendix
AI Analysis in Detail
This chapter lists, for every analysis type of the AI Analysis action, its commands, settings,
output variables, and JSON result format. For the concept and the setup procedure see
Part 4, AI Analysis; for registering analyzers and their
mode and detection FPS see Part 6, AI Analyzer. All
analysis runs on the controller; video and audio are never sent outside.
Common Items
Analyzer and Type
Select the analyzer to use in AI Analyzer. Analyzers can be added,
edited, and deleted directly in the selection popup, and also managed in
Settings > AI Analyzer. The analysis type is determined by the
analyzer's AI Model, so the action dialog shows only the commands and
fields of the model that the selected analyzer uses.
The Type field decides what this action does to the analyzer.
| Type | Description |
| Add Analysis | Adds an analysis to the AI analyzer. A single analyzer can hold several analyses. The command, setting, and output fields below are shown only for this type. |
| Clear Analyses / Clear All Analyses | Removes every analysis registered on this analyzer. Shown as Clear Analyses for video analyzers and Clear All Analyses for audio analyzers. |
| Start Analysis | Starts audio capture and analysis. Audio analyzers only. |
| Stop Analysis | Stops audio capture and analysis. Audio analyzers only. |
Caution
Video analysis runs only after the camera is started from
Action > Camera, and audio analysis runs only after the
Start Analysis command follows Add Analysis. If
either step is missing, the output variables stay empty.
Command
The Command differs per analysis type and decides what is computed and
which variables receive it. Choosing a command shows only the settings and outputs that the
command needs. The commands of each type are listed in the tables of the sections below.
Common Settings
These fields have the same meaning across several types. Where a range or default differs per type, it is noted in that type's section.
| Setting | Description |
| Confidence (%) | Only detections at or above this value are used. The range is 0 to 100; raising it keeps only certain results, lowering it misses fewer. The default is 50 for most types and 70 for face detection. Fields labeled Threshold, Min Confidence, or Detection Threshold have the same meaning. |
| Detection Zone | Drag on the camera preview to define a rectangular area. Without one, the whole frame is analyzed; with one, only results inside it are used. |
| Select By | Decides which detection is written to the variables when several are found. Object Detection and Face Detection offer Highest Confidence, By Confidence Index (0-based), Largest Size, and Closest From Ref. Point; Color Tracking offers Largest, Closest, and By Index; QR / Barcode offers Largest Area, Nearest, and By Index; Edge Impulse Object Detection offers Best Confidence, Index, Largest, and Closest. |
| Index (0-based) | The position to pick when Select By is index-based. Object Detection and Face Detection order by confidence (0, 1, 2… from the highest); Color Tracking and QR / Barcode order by area (0, 1, 2… from the largest); Edge Impulse Object Detection uses the model's detection order. The range is 0 to 15 for objects and faces and 0 to 100 for colors. |
| Reference Point | The reference coordinates, set by clicking on the camera preview, used when Select By is the closest detection. |
Output Variables and Value Conventions
A variable bound to an output field is updated continuously while the analysis runs. Unbound
outputs are not written, so bind variables only to the fields you need. The variable type of
each field is fixed.
- Digital: true/false fields such as Detected, Matched, and Is Match.
- Number: counts, coordinates, width and height, confidence, probability, angle, and distance. Coordinates and sizes are in pixels of the camera frame, and a center point is the center of the detection box. Confidence and probability are 0 to 100 (OCR is the exception, see its section).
- Text: class names, recognized names, data, and JSON output.
- Color: the average color of Color Tracking.
The value written when nothing is detected is defined per type: text becomes none or an
empty string, and numbers become -1 or 0, as listed in each section's output table.
The Output (JSON Text) field writes the whole detection result as a JSON string to a text
variable. Types that return several results use an array, which is empty ([]) when
nothing is detected. The key names and samples are given under JSON Output in each section;
parse the string in block coding to pick out values.
Object Detection
Detects 80 kinds of common objects such as people and vehicles in real time from the camera video. The classes are the 80 of the COCO dataset, and several can be selected together.
Commands
| Command | Description |
| Object Detected? | Checks whether any of the selected classes is detected. |
| Count Objects | Writes the total number of objects of the selected classes. |
| Get Object | Writes the class name, confidence, center point, and width and height of one object of the selected classes, chosen by Select By. |
| Get All Objects | Writes every detected object as JSON. |
When a Detection Zone is set, every command considers only the objects inside it.
Settings
| Setting | Description |
| Object | The kinds of object to detect. Several of the 80 classes can be selected; the default is Person. |
| Confidence (%) | 0 to 100, default 50. |
| Detection Zone | See Common Items. |
| Select By | Highest Confidence (default), By Confidence Index (0-based), Largest Size, Closest From Ref. Point. |
| Index (0-based) | 0 to 15, default 0. Ordered from the highest confidence. |
| Reference Point | See Common Items. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when an object is detected. |
| Count | Number | Number of detected objects. |
| Object | Text | Class name of the selected object (English, e.g. person). |
| Confidence (%) | Number | Confidence of the selected object, 0 to 100. |
| Center Point X / Center Point Y | Number | Pixel coordinates of the selected object's center. |
| Width (Pixels) / Height (Pixels) | Number | Size of the selected object's box. |
| Output (JSON Text) | Text | Array of every detected object. |
JSON Output
An array with, per object, class_name, confidence (0 to 100), the two box corners in bbox, the center, width, and height.
[
{
"class_name": "person",
"confidence": 85,
"bbox": { "x1": 100, "y1": 50, "x2": 300, "y2": 400 },
"center": { "x": 200, "y": 225 },
"width": 200,
"height": 350
},
{
"class_name": "car",
"confidence": 72,
"bbox": { "x1": 400, "y1": 200, "x2": 550, "y2": 300 },
"center": { "x": 475, "y": 250 },
"width": 150,
"height": 100
}
]
The 80 detectable objects
The screen shows the display name; variables and JSON hold the class name in parentheses.
Person (person), Bicycle (bicycle), Car (car), Motorcycle (motorcycle), Airplane (airplane), Bus (bus), Train (train), Truck (truck), Boat (boat), Traffic Light (traffic light), Fire Hydrant (fire hydrant), Stop Sign (stop sign), Parking Meter (parking meter), Bench (bench), Bird (bird), Cat (cat), Dog (dog), Horse (horse), Sheep (sheep), Cow (cow), Elephant (elephant), Bear (bear), Zebra (zebra), Giraffe (giraffe), Backpack (backpack), Umbrella (umbrella), Handbag (handbag), Tie (tie), Suitcase (suitcase), Frisbee (frisbee), Skis (skis), Snowboard (snowboard), Sports Ball (sports ball), Kite (kite), Baseball Bat (baseball bat), Baseball Glove (baseball glove), Skateboard (skateboard), Surfboard (surfboard), Tennis Racket (tennis racket), Bottle (bottle), Wine Glass (wine glass), Cup (cup), Fork (fork), Knife (knife), Spoon (spoon), Bowl (bowl), Banana (banana), Apple (apple), Sandwich (sandwich), Orange (orange), Broccoli (broccoli), Carrot (carrot), Hot Dog (hot dog), Pizza (pizza), Donut (donut), Cake (cake), Chair (chair), Couch (couch), Potted Plant (potted plant), Bed (bed), Dining Table (dining table), Toilet (toilet), TV (tv), Laptop (laptop), Mouse (mouse), Remote (remote), Keyboard (keyboard), Cell Phone (cell phone), Microwave (microwave), Oven (oven), Toaster (toaster), Sink (sink), Refrigerator (refrigerator), Book (book), Clock (clock), Vase (vase), Scissors (scissors), Teddy Bear (teddy bear), Hair Drier (hair drier), Toothbrush (toothbrush)
Face Detection
Detects human faces in real time from the camera video and tracks their positions. Includes emotion recognition and age and gender estimation.
Commands
| Command | Description |
| Face Detected? | Checks whether a face is detected. |
| Count Faces | Writes the number of detected faces. |
| Get Face | Writes the confidence, center point, and width and height of one face chosen by Select By. |
| Get All Faces | Writes every detected face as JSON. |
| Recognize Emotion | When a face is detected, analyzes the expression and classifies it into one of 8 emotions. |
| Estimate Age/Gender | Estimates the age and gender of the detected face. |
When a Detection Zone is set, only faces inside it are considered.
Settings
| Setting | Description |
| Confidence (%) | Default 70. |
| Detection Zone | See Common Items. |
| Select By | Highest Confidence (default), By Confidence Index (0-based), Largest Size, Closest From Ref. Point. |
| Index (0-based) | 0 to 15, default 0. |
| Reference Point | See Common Items. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when a face is detected. |
| Count | Number | Number of detected faces. |
| Confidence (%) | Number | Confidence of the selected face, 0 to 100. |
| Center Point X / Center Point Y | Number | Pixel coordinates of the selected face's center. |
| Width (Pixels) / Height (Pixels) | Number | Size of the selected face's box. |
| Output (JSON Text) | Text | Array of every detected face. |
| Emotion | Text | English name of the recognized emotion; none when not detected. |
| Emotion ID | Number | Emotion number 0 to 7; -1 when not detected. |
| Emotion Confidence (%) | Number | 0 to 100; -1 when not detected. |
| Age | Number | Estimated age 0 to 100; -1 when not detected. |
| Gender | Text | male or female; none when not detected. |
| Gender Confidence (%) | Number | 0 to 100; -1 when not detected. |
JSON Output
Each face carries score (confidence, 0 to 100), bbox, center, width, and height. The emotion keys (emotion, emotion_id, emotion_confidence) are present only while Recognize Emotion runs, and the age and gender keys (age, gender, gender_confidence) only while Estimate Age/Gender runs. Note that the confidence key is score, unlike confidence in Object Detection.
[
{
"score": 95,
"bbox": { "x1": 120, "y1": 80, "x2": 250, "y2": 300 },
"center": { "x": 185, "y": 190 },
"width": 130,
"height": 220,
"emotion": "happy",
"emotion_id": 4,
"emotion_confidence": 87,
"age": 28,
"gender": "male",
"gender_confidence": 92
}
]
The 8 emotions and their IDs
| ID | Emotion | Meaning |
| 0 | angry | Anger |
| 1 | contempt | Contempt |
| 2 | disgust | Disgust |
| 3 | fear | Fear |
| 4 | happy | Happiness |
| 5 | neutral | Neutral |
| 6 | sad | Sadness |
| 7 | surprise | Surprise |
Face Recognition
An extension of Face Detection that compares detected faces against enrolled faces to identify who they are. Enrolling and deleting faces is also done through this action's commands.
Commands
| Command | Description |
| Recognize Face | Identifies which enrolled face the current face is and writes the name and similarity. |
| Get Enrolled Count | Writes the total number of enrolled faces. |
| Get Enrolled List | Writes the enrolled face names as a comma-separated string. |
| Enroll Face | Enrolls the face currently on screen under Face Name. |
| Delete Face | Deletes the enrolled face whose name equals Face Name. |
| Clear All Faces | Deletes every enrolled face. |
Settings
| Setting | Description |
| Threshold (%) | Similarity threshold. 0 to 100, default 70. A face counts as matched only when the similarity is at or above this value. |
| Min Face Area (px²) | Minimum pixel area of a face to recognize. Faces smaller (farther) than this are excluded. 0 means no limit. The input is normalized to 112×112, so very small faces lose accuracy when upscaled. The default 6400 (80×80) is a balanced value; raise it to 12544 (112×112) for close-range, high-precision recognition only. |
| Face Name | Name of the face to enroll or delete. Example: john, Alice. |
The face store is chosen by the analyzer setting DB Name (default
default); analyzers with the same name share data. Anti-Spoofing,
which blocks faces shown on photos or screens, is also an analyzer setting
(Part 6, AI Analyzer).
Output Variables
| Field | Type | Value |
| Face Detected | Digital | True when a face is detected. |
| Matched | Digital | True when an enrolled face matches at or above the threshold. |
| Recognized Name | Text | Name of the matched face. An empty string when there is no face; Unknown when a face is present but matches no enrolled face. |
| Similarity (%) | Number | Similarity to the closest enrolled face, 0 to 100. |
| Enrolled Count | Number | Number of enrolled faces. |
| Enrolled List | Text | Enrolled names, comma-separated. |
This type has no JSON output.
Pose Estimation
Estimates a person's pose in real time from the camera video and provides the coordinates of 17 key points (joints). Distances between key points and joint angles can be computed directly.
Commands
| Command | Description |
| Person Detected? | Checks whether a person is detected. |
| Get Keypoint | Writes the coordinates and confidence of the specified key point (e.g. wrist, knee). |
| Distance Between Keypoints | Writes the pixel distance between two key points. |
| Joint Angle | Computes and writes the angle of a joint formed by three key points (e.g. the elbow angle). |
| Keypoint Exists in Zone | Checks whether the specified key point lies inside the Detection Zone. |
| Get All Keypoints | Writes the coordinates of all 17 key points as JSON. |
Settings
| Setting | Description |
| Key Point | The key point whose coordinates are read or whose presence in the zone is checked. One of 17. |
| Key Point 1 / Key Point 2 | The two key points whose distance is computed. |
| Joint | The joint whose angle is computed: Left Elbow, Right Elbow, Left Shoulder, Right Shoulder, Left Knee, Right Knee, Left Hip, Right Hip, Spine Tilt, or Neck Tilt. |
| Detection Zone | See Common Items. |
Output Variables
| Field | Type | Value |
| Detected | Digital | Result of Person Detected?, Get Keypoint, and Keypoint Exists in Zone. |
| Key Point X / Key Point Y | Number | Pixel coordinates of the specified key point. |
| Confidence (%) | Number | Confidence of the specified key point, 0 to 100. |
| Detected | Digital | For Distance Between Keypoints and Joint Angle: true when both key points (the joint) were detected and the value was computed. |
| Joint Angle (°) | Number | Angle of the specified joint. |
| Distance (pixels) | Number | Distance between the two key points. |
| Output (JSON Text) | Text | Array of the 17 key points. |
JSON Output
An array with, per key point, id, name, x, y, and confidence (0 to 100). The sample shows only a few entries.
[
{ "id": 0, "name": "nose", "x": 320, "y": 180, "confidence": 95 },
{ "id": 5, "name": "left_shoulder", "x": 280, "y": 260, "confidence": 91 },
{ "id": 6, "name": "right_shoulder", "x": 360, "y": 258, "confidence": 93 },
{ "id": 15, "name": "left_ankle", "x": 270, "y": 470, "confidence": 78 }
]
The 17 key points
| ID | JSON name | Shown as |
| 0 | nose | Nose |
| 1 | left_eye | Left Eye |
| 2 | right_eye | Right Eye |
| 3 | left_ear | Left Ear |
| 4 | right_ear | Right Ear |
| 5 | left_shoulder | Left Shoulder |
| 6 | right_shoulder | Right Shoulder |
| 7 | left_elbow | Left Elbow |
| 8 | right_elbow | Right Elbow |
| 9 | left_wrist | Left Wrist |
| 10 | right_wrist | Right Wrist |
| 11 | left_hip | Left Hip |
| 12 | right_hip | Right Hip |
| 13 | left_knee | Left Knee |
| 14 | right_knee | Right Knee |
| 15 | left_ankle | Left Ankle |
| 16 | right_ankle | Right Ankle |
Hand Tracking
Tracks hands in real time from the camera video and provides the coordinates of 21 key points (finger joints). Includes left/right handedness, finger extension state, and gesture recognition.
Commands
| Command | Description |
| Hand Detected? | Checks whether a hand is detected. |
| Get Keypoint | Writes the X, Y, and Z coordinates of the specified key point. |
| Get Handedness | Writes whether the detected hand is the left or the right hand. |
| Get Finger State | Writes whether the specified finger is extended or bent. |
| Hand Exists in Zone | Checks whether a hand lies inside the Detection Zone. |
| Keypoint Exists in Zone | Checks whether the specified key point lies inside the Detection Zone. |
| Distance Between Keypoints | Writes the pixel distance between two key points. |
| Get All Keypoints | Writes the coordinates of all 21 key points as JSON. |
| Recognize Gesture | Classifies the hand gesture (e.g. fist, victory, thumbs-up) from the layout of the 21 key points and writes its name and ID. |
Settings
| Setting | Description |
| Confidence (%) | 0 to 100, default 50. |
| Key Point | The key point whose coordinates are read or whose presence in the zone is checked. One of 21. |
| Key Point 1 / Key Point 2 | The two key points whose distance is computed. |
| Finger | The finger whose state is checked: Thumb, Index, Middle, Ring, or Pinky. |
| Detection Zone | See Common Items. |
Output Variables
| Field | Type | Value |
| Detected | Digital | Result of Hand Detected?, Get Keypoint, Hand Exists in Zone, and Keypoint Exists in Zone. |
| Key Point X / Key Point Y / Key Point Z | Number | Coordinates of the specified key point. X and Y are pixel coordinates; Z is the depth value estimated by the model. |
| Handedness Value | Number | 0 for the left hand, 1 for the right hand, -1 when it cannot be determined. |
| Handedness Label | Text | left or right; an empty string when it cannot be determined. |
| Distance (pixels) | Number | Distance between the two key points. |
| Detected | Digital | For Distance Between Keypoints: true when both key points were detected and the distance was computed. |
| Finger State | Number | 0 for bent, 1 for extended, -1 when it cannot be detected. |
| Output (JSON Text) | Text | Array of the 21 key points. |
| Gesture Name | Text | English name of the recognized gesture; none when not detected. |
| Gesture ID | Number | Gesture number; -1 when not detected. |
JSON Output
An array with, per key point, id, name, x, y, and z. The sample shows only a few entries.
[
{ "id": 0, "name": "wrist", "x": 320, "y": 350, "z": 0 },
{ "id": 4, "name": "thumb_tip", "x": 280, "y": 310, "z": -15 },
{ "id": 8, "name": "index_tip", "x": 310, "y": 260, "z": -25 },
{ "id": 12, "name": "middle_tip", "x": 330, "y": 255, "z": -22 }
]
The 21 key points
| ID | JSON name | Shown as |
| 0 | wrist | Wrist |
| 1 | thumb_cmc | Thumb Wrist Joint |
| 2 | thumb_mcp | Thumb Knuckle |
| 3 | thumb_ip | Thumb Joint |
| 4 | thumb_tip | Thumb Tip |
| 5 | index_mcp | Index Knuckle |
| 6 | index_pip | Index 1st Joint |
| 7 | index_dip | Index 2nd Joint |
| 8 | index_tip | Index Tip |
| 9 | middle_mcp | Middle Knuckle |
| 10 | middle_pip | Middle 1st Joint |
| 11 | middle_dip | Middle 2nd Joint |
| 12 | middle_tip | Middle Tip |
| 13 | ring_mcp | Ring Knuckle |
| 14 | ring_pip | Ring 1st Joint |
| 15 | ring_dip | Ring 2nd Joint |
| 16 | ring_tip | Ring Tip |
| 17 | pinky_mcp | Pinky Knuckle |
| 18 | pinky_pip | Pinky 1st Joint |
| 19 | pinky_dip | Pinky 2nd Joint |
| 20 | pinky_tip | Pinky Tip |
The 33 gestures and their IDs
Gesture ID is the number in this table; Gesture Name is the English name.
| ID | Name | Meaning |
| 0 | call | Phone gesture (thumb and pinky extended) |
| 1 | dislike | Thumbs down |
| 2 | fist | Closed fist |
| 3 | four | Number 4 (thumb folded) |
| 4 | grabbing | Grabbing |
| 5 | grip | Grip |
| 6 | hand_heart | Finger heart (thumb and index) |
| 7 | hand_heart2 | Finger heart (alternate form) |
| 8 | holy | Prayer hands |
| 9 | like | Thumbs up |
| 10 | little_finger | Pinky extended |
| 11 | middle_finger | Middle finger extended |
| 12 | mute | Shush (index finger to lips) |
| 13 | ok | OK sign (thumb and index circle) |
| 14 | one | Number 1 (index extended) |
| 15 | palm | Open palm |
| 16 | peace | Peace / victory |
| 17 | peace_inverted | Inverted peace |
| 18 | point | Pointing (index finger) |
| 19 | rock | Rock (index and pinky) |
| 20 | stop | Stop (palm forward) |
| 21 | stop_inverted | Inverted stop |
| 22 | take_picture | Take-picture gesture |
| 23 | three | Number 3 |
| 24 | three2 | Number 3 (alternate form) |
| 25 | three3 | Number 3 (third form) |
| 26 | three_gun | Finger gun (three fingers) |
| 27 | thumb_index | Thumb and index extended |
| 28 | thumb_index2 | Thumb and index extended (alternate form) |
| 29 | timeout | Timeout (T shape) |
| 30 | two_up | Number 2 (two fingers up) |
| 31 | two_up_inverted | Inverted number 2 |
| 32 | xsign | X sign (two index fingers crossed) |
Color Tracking
Finds regions of a specified color range in real time from the camera video and provides their position, size, area, and average color.
Commands
| Command | Description |
| Color Detected? | Checks whether the specified color range is detected. |
| Count Colors | Writes the number of detected color regions. |
| Get Color Info | Writes the center point, width and height, area, and average color of one color region chosen by Select By. |
| Get All Colors | Writes every detected color region as JSON. |
When a Detection Zone is set, only color regions inside it are considered.
Settings
| Setting | Description |
| Start Color / End Color | The color range to detect. Colors between the two are detected. Defaults are #FF0000 to #FF5555. |
| Min Area (px²) | Regions smaller than this are ignored, which filters out small noise. 1 to 100000, default 300. |
| Detection Zone | See Common Items. |
| Select By | Largest (default), Closest, By Index. |
| Reference Point | See Common Items. |
| Index (0-based) | 0 to 100, default 0. Ordered from the largest area. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when a color region is detected. |
| Count | Number | Number of detected color regions. |
| Center Point X / Center Point Y | Number | Pixel coordinates of the selected region's center; 0 when not detected. |
| Width (Pixels) / Height (Pixels) | Number | Size of the selected region; 0 when not detected. |
| Area | Number | Area of the selected region in px²; 0 when not detected. |
| Avg Color | Color | Average color of the selected region. |
| Output (JSON Text) | Text | Array of every detected color region. |
JSON Output
An array with, per region, the center x and y, width, height, area, and the average color as a hex string.
[
{ "x": 320, "y": 240, "width": 50, "height": 60, "area": 3000, "color": "#FF5733" },
{ "x": 500, "y": 300, "width": 40, "height": 45, "area": 1800, "color": "#33FF57" }
]
QR / Barcode
Detects and decodes QR codes and several one-dimensional barcode formats in real time.
Commands
| Command | Description |
| Scan | Reads one code and writes its data, code type, position, and size. When several codes are detected, one is chosen by Select By. |
| Scan All | Writes every detected code as a JSON array. |
| Has Code? | Checks whether a code is detected. |
| Count Codes | Writes the number of detected codes. |
Settings
| Setting | Description |
| Code Type | Detects only codes of one type. The default All detects every type. The available types are listed below. |
| Select By | Largest Area (default), Nearest, By Index. |
| Index (from 0) | The position to pick when Select By is By Index. Ordered from the largest area; default 0. |
| Reference Point | See Common Items. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when a code is detected. |
| Data | Text | The decoded string. |
| Code Type | Text | The code type name reported by the decoder (e.g. QR-Code, EAN-13). |
| Center X / Center Y | Number | Pixel coordinates of the code's center. |
| Width (px) / Height (px) | Number | Size of the code area. |
| Count | Number | Number of detected codes. |
| Output (JSON Text) | Text | Array of every detected code. |
JSON Output
An array with, per code, data, type, the center x and y, width, and height.
[
{ "data": "https://example.com", "type": "QR-Code", "x": 320, "y": 240, "width": 100, "height": 100 },
{ "data": "1234567890123", "type": "EAN-13", "x": 500, "y": 300, "width": 80, "height": 40 }
]
Code Type options
All, QR Code, EAN-8, EAN-13, UPC-A, UPC-E, Code-128, Code-39, Code-93, Codabar, ITF, DataBar, DataBar-Expanded
OCR (Text Recognition)
Detects and reads text in real time from the camera video. The recognition language is chosen in the analyzer setting OCR Language: English (default), Korean, Japanese, Chinese (Simplified), German, French, Spanish, Portuguese, Hindi, Arabic, or Russian. One analyzer recognizes one language.
- Temporal stabilization suppresses results that flicker from frame to frame.
- Detected text is sorted into reading order (top to bottom, left to right).
Commands
| Command | Description |
| Scan | Merges the text items at or above Min Confidence in reading order and writes the result; multiple lines are separated by the newline character \n. When a Recognized Text variable is bound, results from several frames are compared and the variable is updated only when a stable text is established or changes, and it is cleared when the text has been absent for a while. |
| Scan All | Writes every detected text item as a JSON array. |
| Has Text? | Checks whether text is detected. |
| Count Texts | Writes the number of detected text items. |
Settings
| Setting | Description |
| Min Confidence (%) | Only text with confidence at or above this value is included. 0 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when text is detected. |
| Recognized Text | Text | The merged string in reading order; lines separated by \n. |
| Confidence (%) | Number | Average confidence of the included text. Unlike other analysis types, this is written as a fraction between 0 and 1. |
| Count | Number | Number of detected text items. |
| Output (JSON Text) | Text | Array of every detected text item. |
JSON Output
An array with, per item, text, confidence (a fraction 0 to 1), the center x and y, width, and height.
[
{ "text": "Hello", "confidence": 0.95, "x": 100, "y": 50, "width": 80, "height": 25 },
{ "text": "World", "confidence": 0.88, "x": 190, "y": 50, "width": 70, "height": 25 }
]
Line Tracking
For line-following robots. Detects a line on the floor and provides its offset from the frame center, its tilt angle, and its vector, which are used to correct steering.
Commands
| Command | Description |
| Line Info | Writes the detection result, offset, angle, vector, and line width of the detected line. This is the core command for steering control. |
| Has Line? | Checks whether a line is detected. |
| All Info (JSON) | Writes all line tracking information as JSON. |
Settings
| Setting | Description |
| Color Mode | Dark Line on Light BG (default), Light Line on Dark BG, or Custom Color. Custom Color detects lines whose color lies between the Start Color and End Color below. |
| Start Color / End Color | The color range for Custom Color mode. Both default to #FF0000. |
| Sensitivity (%) | 0 to 100, default 50. Higher values detect faint lines but increase false detections; lower values detect only clear lines. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when a line is detected. |
| Offset X (-1~1) | Number | Horizontal position of the line relative to the frame center: -1 (left edge) to 1 (right edge), 0 when centered. Used for left/right correction. |
| Angle (deg) | Number | Tilt of the line, -90 to 90 degrees, 0 when vertical. Used for heading correction. |
| Vector X0 / Y0 / X1 / Y1 | Number | Pixel coordinates of the line vector's start (X0, Y0) and end (X1, Y1). |
| Line Width | Number | Line width as a fraction of the frame width, 0 to 1. |
| Output (JSON Text) | Text | A JSON object with all values. |
When no line is detected, Detected becomes false and the remaining numeric values are written as 0.
JSON Output
A single object, not an array. offset_x and line_width are fractions as described in the table.
{
"detected": true,
"offset_x": 0.15,
"angle": 12.5,
"line_width": 0.06,
"vector_x0": 100,
"vector_y0": 200,
"vector_x1": 540,
"vector_y1": 210
}
Line Crossing Counter
Automatically counts people entering and exiting across a virtual line drawn on the camera preview. Used for entrance footfall and store visitor counting.
Commands
| Command | Description |
| Get Count | Writes the current in, out, and net counts to variables. |
| Reset Count | Resets all counts to 0. |
Settings
| Setting | Description |
| Counting Line | Drawn by dragging on the camera preview. The arrow shown on the line is the IN (entry) direction; crossings in the opposite direction count as OUT (exit). |
| Confidence (%) | 10 to 100, default 50. Only people detected at or above this value are counted. |
Output Variables
| Field | Type | Value |
| In Count | Number | Cumulative crossings in the IN direction. |
| Out Count | Number | Cumulative crossings in the OUT direction. |
| Net Count | Number | Current occupancy, In Count minus Out Count. |
| JSON Result | Text | A JSON object with the three counts and the list of tracked people. |
JSON Output
tracks holds, per tracked person, the track ID, class, center x and y, size w and h, crossing direction (in or out), and whether it has crossed.
{
"in_count": 5,
"out_count": 3,
"current_count": 2,
"tracks": [
{ "id": 1, "class": "person", "x": 320, "y": 240, "w": 80, "h": 200, "direction": "in", "crossed": true }
]
}
Fire Detection
Detects flames in real time from the camera video. A digital variable that turns true when the fire probability exceeds the threshold drives automatic responses such as alarms, notifications, and equipment shutdown.
Commands
| Command | Description |
| Fire Detection | Checks whether fire is detected. |
| Detection Result | Writes the fire probability of the current frame. |
Settings
| Setting | Description |
| Detection Threshold (%) | Probability at or above this value is classified as fire. 10 to 100, default 30. Raise it when there are false alarms, lower it when fires are missed. |
Output Variables
| Field | Type | Value |
| Fire Probability | Number | Fire probability of the current frame, 0 to 100. |
| Fire Detected | Digital | True when the probability is at or above the threshold. |
This type has no JSON output. The analyzer's Mode field runs the same model whichever value is chosen.
Action Recognition
Recognizes what a person is doing from 400 action categories (walking, running, sitting, clapping, cooking, and more). Used for fall detection, exercise form checks, and behavior analysis.
Commands
| Command | Description |
| Specific Action? | Checks whether the target action is detected at or above the threshold. |
| Recognition Result | Writes the name and probability of the most probable action. |
| All Results | Writes every recognition result, ordered by probability, as JSON. |
Settings
| Setting | Description |
| Target Action | The action watched by Specific Action?. One of the 400 actions of the Kinetics-400 dataset; the default is abseiling. |
| Detection Threshold (%) | Probability at or above this value counts as the action. 10 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Action Name | Text | English name of the most probable action. |
| Confidence | Number | Probability of the recognized action, 0 to 100. |
| Action Matched | Digital | True when Specific Action? finds the target action at or above the threshold. |
| JSON Result | Text | Array of every recognition result. |
JSON Output
An array with, per result, label, confidence (0 to 100), and the model class number index, ordered from the highest probability.
[
{ "label": "walking", "confidence": 85, "index": 0 },
{ "label": "running", "confidence": 10, "index": 1 }
]
For the full list of the 400 actions see the official label list of the Kinetics-400 dataset (github.com/cvdfoundation/kinetics-dataset).
Hand Motion Recognition
Recognizes 25 hand motions performed in front of the camera (swipes, thumb up, stop sign, and more). Used for touchless control and game input. Unlike the gesture recognition of Hand Tracking, which classifies a hand shape in a single frame, hand motion recognition classifies movement over time.
Commands
| Command | Description |
| Specific Hand Motion? | Checks whether the target hand motion is detected at or above the threshold. |
| Recognition Result | Writes the name and probability of the most probable hand motion. |
| All Results | Writes every recognition result, ordered by probability, as JSON. |
Settings
| Setting | Description |
| Target Hand Motion | The motion watched by Specific Hand Motion?. One of 25; the default is Swiping Left. |
| Detection Threshold (%) | Probability at or above this value counts as the motion. 10 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Hand Motion Name | Text | English name of the most probable hand motion. |
| Confidence | Number | Probability of the recognized motion, 0 to 100. |
| Hand Motion Matched | Digital | True when Specific Hand Motion? finds the target motion at or above the threshold. |
| JSON Result | Text | Array of every recognition result. |
JSON Output
An array with, per result, label, confidence (0 to 100), and the model class number index, ordered from the highest probability.
[
{ "label": "Swiping Left", "confidence": 85, "index": 0 },
{ "label": "Thumb Up", "confidence": 10, "index": 1 }
]
The 25 hand motions
Swiping Left, Swiping Right, Swiping Down, Swiping Up, Pushing Hand Away, Pulling Hand In, Sliding Two Fingers Left, Sliding Two Fingers Right, Sliding Two Fingers Down, Sliding Two Fingers Up, Pushing Two Fingers Away, Pulling Two Fingers In, Rolling Hand Forward, Rolling Hand Backward, Turning Hand Clockwise, Turning Hand Counterclockwise, Zooming In With Full Hand, Zooming Out With Full Hand, Zooming In With Two Fingers, Zooming Out With Two Fingers, Thumb Up, Thumb Down, Shaking Hand, Stop Sign, Drumming Fingers
License Plate Recognition
Detects vehicle license plates in real time from the camera video and reads their text. The plate model is chosen in the analyzer setting Region: Latin (EU/US/SouthAmerica) (default) or Korea.
- Results from several frames are corrected by majority voting, which automatically fixes single-frame misreads.
- A minimum confidence filter removes noisy results.
Commands
| Command | Description |
| Scan | Writes the text, region, and confidence of the top-confidence plate. |
| Scan All | Writes every detected plate as a JSON array. |
| Has Plate? | Checks whether a plate is detected. |
| Count Plates | Writes the number of detected plates. |
Settings
| Setting | Description |
| Min Confidence (%) | Only plates with confidence at or above this value are included. 0 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when a plate is detected. |
| Plate Text | Text | The recognized plate string. |
| Detected Region | Text | The country of the plate as classified by the model (e.g. Argentina, Germany, Korea). |
| Confidence (%) | Number | Average text recognition confidence, 0 to 100. |
| Count | Number | Number of detected plates. |
| Output (JSON Text) | Text | Array of every detected plate. |
JSON Output
An array with, per plate, text, region, confidence (text recognition confidence, 0 to 100), the plate detection stage score det_score, the center x and y, width, and height.
[
{ "text": "ABC1234", "region": "Argentina", "confidence": 92, "det_score": 0.97, "x": 320, "y": 240, "width": 120, "height": 40 }
]
Teachable Machine Classification
Classifies the camera video with an image classification model trained in Google Teachable Machine. The model is uploaded in the analyzer settings. Model Path (.tflite) accepts only 224×224 RGB models in TensorFlow Lite format, exported as the Floating Point or Quantized type. Label Path is the labels.txt exported with the model, with one class name per line in order. For the procedure to create a model see Appendix B, External Services.
Commands
| Command | Description |
| Get Classification Result | Writes the class name and confidence of the most probable class. |
| Is Specific Class? | Checks whether the result matches the Target Class with confidence at or above the threshold. |
| Get All Results | Writes the probability of every class as JSON. |
Settings
| Setting | Description |
| Target Class | The class (label) name compared by Is Specific Class?. It must exactly match a name in labels.txt. Example: cat, defect. |
| Threshold (%) | A result counts as matched only when its confidence is at or above this value. 0 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Class Name | Text | Name of the most probable class. |
| Confidence (%) | Number | Confidence of that class, 0 to 100. |
| Is Match | Digital | True when the class equals the Target Class at or above the threshold. |
| Output (JSON Text) | Text | Array of the probability of every class. |
JSON Output
An array with, per class, label and confidence (0 to 100), ordered from the highest probability.
[
{ "label": "apple", "confidence": 92 },
{ "label": "banana", "confidence": 6 },
{ "label": "orange", "confidence": 2 }
]
Edge Impulse Classification
Classifies images or sound with a classification model trained in Edge Impulse. Video and audio analyzers use the same commands and fields; the model is uploaded as the .eim file exported from Edge Impulse Studio in the analyzer setting Model Path (.eim). For the procedure to create a model see Appendix B, External Services.
Commands
| Command | Description |
| Get Classification Result | Writes the class name and confidence of the most probable class. |
| Is Specific Class? | Checks whether the result matches the Target Class with confidence at or above the threshold. |
| Get All Results | Writes the probability of every class as JSON. |
| Is Anomaly? | Checks whether the model judged the input as an anomaly. Meaningful only for models that include anomaly detection. |
| Get Anomaly Score | Writes the anomaly detection score. |
Settings
| Setting | Description |
| Target Class | The class (label) name compared by Is Specific Class?. It must exactly match a label in the model. Example: cat, defect. |
| Threshold (%) | A result counts as matched only when its confidence is at or above this value. 0 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Class Name | Text | Name of the most probable class. |
| Confidence (%) | Number | Confidence of that class, 0 to 100. |
| Is Match | Digital | True when the class equals the Target Class at or above the threshold. |
| Is Anomaly | Digital | True when judged as an anomaly. |
| Anomaly Score | Number | The anomaly detection score. |
| Output (JSON Text) | Text | Array of the probability of every class. |
JSON Output
An array with, per class, label and confidence (0 to 100), ordered from the highest probability.
[
{ "label": "normal", "confidence": 89 },
{ "label": "defective", "confidence": 11 }
]
Edge Impulse Object Detection
Detects objects in the camera video with an object detection model trained in Edge Impulse. The model is uploaded in the analyzer setting Model Path (.eim).
Commands
| Command | Description |
| Has Object | Checks whether an object is detected: the label typed in Object if set, any object if it is empty. |
| Count Objects | Writes the number of objects with that label (or all objects). |
| Get Object Info | Writes the class name, confidence, center point, and width and height of one object chosen by Select By. |
| Get All Objects Info | Writes every detected object as JSON. |
| Is Anomaly | Checks whether the model judged the input as an anomaly. |
| Get Anomaly Score | Writes the anomaly detection score. |
When a Detection Zone is set, only objects inside it are considered.
Settings
| Setting | Description |
| Object | The class (label) name to look for, typed directly. It must exactly match a label in the model; leave it empty to target all objects. |
| Confidence (%) | 0 to 100, default 50. |
| Detection Zone | See Common Items. |
| Select By | Best Confidence (default), Index, Largest, Closest. |
| Index (0-based) | The position to pick when Select By is Index. Uses the model's detection order; default 0. |
| Reference Point | See Common Items. |
Output Variables
| Field | Type | Value |
| Detected | Digital | True when an object is detected. |
| Count | Number | Number of detected objects. |
| Class Name | Text | Label of the selected object; an empty string when not detected. |
| Confidence (%) | Number | Confidence of the selected object, 0 to 100; -1 when not detected. |
| Center Point X / Center Point Y | Number | Pixel coordinates of the selected object's center; -1 when not detected. |
| Width (Pixels) / Height (Pixels) | Number | Size of the selected object's box. |
| Is Anomaly | Digital | True when judged as an anomaly. |
| Anomaly Score | Number | The anomaly detection score. |
| Output (JSON Text) | Text | Array of every detected object. |
JSON Output
A flat array with, per object, label, confidence (0 to 100), the box's top-left x and y, width, and height. Unlike the built-in Object Detection there are no nested bbox and center objects, and x and y are not the center.
[
{ "label": "person", "confidence": 87, "x": 100, "y": 50, "width": 200, "height": 350 },
{ "label": "car", "confidence": 75, "x": 400, "y": 200, "width": 150, "height": 100 }
]
Edge Impulse Visual Anomaly
Judges normal versus abnormal states with a visual anomaly detection model trained in Edge Impulse. The frame is divided into a grid, each cell receives an anomaly score, and the mean and maximum over the grid are provided. The model is uploaded in the analyzer setting Model Path (.eim).
Commands
| Command | Description |
| Get Result | Writes the mean score, max score, and anomaly verdict to variables. |
| Is Anomaly? | Checks whether the result is an anomaly by the threshold. |
| Get Grid Info | Writes the anomaly score of every grid cell as JSON. |
Settings
| Setting | Description |
| Anomaly Threshold (%) | Scores above this value are judged as anomalies. 0 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Mean Score | Number | Mean anomaly score over all grid cells, 0 to 100. |
| Max Score | Number | Highest anomaly score among the grid cells, 0 to 100. |
| Is Anomaly | Digital | True when judged as an anomaly. |
| Output (JSON Text) | Text | A JSON object with the mean, max, and grid. |
JSON Output
A single object, not an array. Each cell in grid has its position x and y, size width and height, and anomaly score value (0 to 100).
{
"mean": 45,
"max": 78,
"grid": [
{ "x": 0, "y": 0, "width": 32, "height": 32, "value": 12 },
{ "x": 32, "y": 0, "width": 32, "height": 32, "value": 67 },
{ "x": 64, "y": 0, "width": 32, "height": 32, "value": 78 }
]
}
Sound Classification
A type of the audio analyzer. Classifies the sound from the microphone input (or the speaker output) into one of the 521 AudioSet categories learned by Google's YAMNet model (speech, music, dog barking, alarms, glass breaking, screaming, car horns, and more). Used for abnormal sound detection and environmental sound analysis. The input is chosen in the analyzer setting Audio Source: Microphone (Input) or Speaker (Output).
Note
Audio analysis writes results only after the Start Analysis command is run
following Add Analysis.
Commands
| Command | Description |
| Specific Sound? | Checks whether the target sound is detected at or above the threshold. |
| Detection Result | Writes the name and probability of the most probable sound. |
| All Results | Writes every recognition result, ordered by probability, as JSON. |
Settings
| Setting | Description |
| Target Sound | The sound watched by Specific Sound?. One of the 521 categories (English names); the default is Speech. |
| Detection Threshold (%) | Probability at or above this value counts as the sound. 10 to 100, default 50. |
Output Variables
| Field | Type | Value |
| Sound Name | Text | English name of the most probable sound. |
| Sound Probability | Number | Probability of the recognized sound, 0 to 100. |
| Sound Detected | Digital | True when Specific Sound? finds the target sound at or above the threshold. |
| JSON Result | Text | Array of every recognition result. |
JSON Output
An array with, per result, label and confidence (0 to 100), ordered from the highest probability. Unlike Action Recognition and Hand Motion Recognition there is no index key.
[
{ "label": "Speech", "confidence": 85 },
{ "label": "Music", "confidence": 10 }
]
For the full list of the 521 categories see YAMNet's official class map (yamnet_class_map.csv).
Edge Impulse Classification (Audio)
Choosing Edge Impulse Classification as the audio analyzer's AI Model recognizes and classifies sound with an audio classification model trained in Edge Impulse. The commands, settings, output variables, and JSON format are the same as Edge Impulse Classification.
- The model is uploaded in the analyzer setting Model Path (.eim), and the input is chosen in Audio Source as the microphone or the speaker.
- Results are written only after the Start Analysis command is run following Add Analysis.