Available perception
| Component | Backend | Output and limits |
|---|---|---|
| Objects/people | YOLO26 Nano, COCO 80 classes | Boxes, classes, confidence, centers; class support does not include arbitrary signs/roads |
| Tracking | ByteTrack | Stream-local IDs and temporal image-plane motion; no tracking across unrelated photos |
| OCR | EasyOCR | Text polygons and confidence, English and configurable Arabic; low-score text filtered, empty results do not prove no text |
| Pose/activity | YOLO26 Nano Pose | 17 keypoints; conservative static arm-raised geometry; walking/running/waving remain unknown |
| Faces | MediaPipe Face Landmarker | Observable blendshapes and tentative smiling-looking result for sufficiently large faces; no identity or inner-emotion inference |
| Depth | Depth Anything V2 Small | Relative heatmap; no metric distance |
| Walkable candidates | SegFormer B0 / Cityscapes | Sidewalk mask with confidence threshold, uncertain road regions, detected-object boxes marked as obstructions; experimental, not navigation certification |
Face outputs are independent of person boxes and are not identity-linked. Small/occluded faces often produce no landmarks or unknown expression. No general-purpose temporal action classifier or SLAM trajectory estimator is implemented. Segmentation does not fuse depth into a calibrated traversability model; never drive a physical robot solely from these outputs. The default detector is optimized for speed; the full CPU pipeline is substantially slower.
Source captured: 2026-10-11