Qentrah
Projects
Guides

Available perception

View as Markdown
ComponentBackendOutput and limits
Objects/peopleYOLO26 Nano, COCO 80 classesBoxes, classes, confidence, centers; class support does not include arbitrary signs/roads
TrackingByteTrackStream-local IDs and temporal image-plane motion; no tracking across unrelated photos
OCREasyOCRText polygons and confidence, English and configurable Arabic; low-score text filtered, empty results do not prove no text
Pose/activityYOLO26 Nano Pose17 keypoints; conservative static arm-raised geometry; walking/running/waving remain unknown
FacesMediaPipe Face LandmarkerObservable blendshapes and tentative smiling-looking result for sufficiently large faces; no identity or inner-emotion inference
DepthDepth Anything V2 SmallRelative heatmap; no metric distance
Walkable candidatesSegFormer B0 / CityscapesSidewalk mask with confidence threshold, uncertain road regions, detected-object boxes marked as obstructions; experimental, not navigation certification

Face outputs are independent of person boxes and are not identity-linked. Small/occluded faces often produce no landmarks or unknown expression. No general-purpose temporal action classifier or SLAM trajectory estimator is implemented. Segmentation does not fuse depth into a calibrated traversability model; never drive a physical robot solely from these outputs. The default detector is optimized for speed; the full CPU pipeline is substantially slower.

Source captured: 2026-10-11