Use and adapt ai-vision-engine
Use and adapt ai-vision-engine: Analyze images and video locally with object detection, tracking, OCR, pose estimation, and relative depth.
Use the documented interfaces and source layout for ai-vision-engine. The sections below retain the README’s examples and configuration details.
Images, video, and cameras
uv run ai-vision benchmark --source path/to/jpeg-directory --output outputs/my-images
uv run ai-vision video path/to/video.mp4 --output outputs/video
uv run ai-vision camera --device 0 --display
uv run ai-vision --compute cpu camera --device 1 --display --max-frames 300Video outputs contain a tracked MP4 and frame-by-frame JSONL. Camera access is explicitly initiated by the camera command. --display opens the OpenCV preview; Q exits. Without it, capture is headless. Set --max-frames to bound recording. ByteTrack assigns temporary stream-local IDs; motion history estimates speed in pixels/second from consecutive frames. Newly/no-longer visible events are observations, not proof of physical entry/exit. Camera movement and ID switches affect these estimates. Independent images never share tracking IDs.
Device selection prefers CUDA, then Apple MPS, then CPU. --compute cpu overrides it. YOLO MPS unsupported-operator errors retry on CPU with a warning. Optional OCR and Transformers modules deliberately run on CPU for compatibility.
Reproduce the 100-image experiment
uv run ai-vision collect --count 100 --seed 20261010
uv run ai-vision benchmark --degrade --output outputs/pretrained
uv run ai-vision validate --output runs/pretrained
uv run ai-vision train --epochs 3 --output runs/finetune
uv run ai-vision --model runs/finetune/run/weights/best.pt validate --output runs/finetuned-validation
uv run ai-vision --ocr --pose --face --depth --segmentation benchmark --output outputs/full-perception
uv run ai-vision film --source outputs/full-perception --seconds 50 --output outputs/detections.mp4Collection downloads COCO's annotation archive (about 242 MiB), then 100 individual images rather than the entire image dataset. It randomly samples images with vehicle, bicycle, traffic-light, or stop-sign annotations and CC BY, CC BY-SA, or government-work metadata. This filter creates diverse street-context examples but also some indoor/transport contexts; it does not guarantee all images are streets.
The originals remain unchanged. Each manifest entry records original Flickr URL, download URL, source license, resolution, timestamp, and SHA256. Respect attribution and share-alike terms for image derivatives. ATTRIBUTION.md provides per-image source links and credit identifiers from COCO metadata.
Real COCO annotations are converted to YOLO labels, excluding crowd regions. The seeded order assigns 80 training and 20 validation images. The short run freezes the first 10 layers, uses AdamW, and retains the best validation checkpoint. This is fine-tuning a pretrained model, not training from scratch. The pretrained baseline performed better, so it remains the default and powers the film. The fine-tuned checkpoint is included as an experiment in the release.
COCO validation data contributed to development of the upstream pretrained model. The holdout is separate from this fine-tuning run, but it is not an independent generalization benchmark. Only detection has labeled ground truth; OCR, pose, face, depth, and segmentation outputs are qualitative. Confidence is not accuracy.
Available perception
| Component | Backend | Output and limits |
|---|---|---|
| Objects/people | YOLO26 Nano, COCO 80 classes | Boxes, classes, confidence, centers; class support does not include arbitrary signs/roads |
| Tracking | ByteTrack | Stream-local IDs and temporal image-plane motion; no tracking across unrelated photos |
| OCR | EasyOCR | Text polygons and confidence, English and configurable Arabic; low-score text filtered, empty results do not prove no text |
| Pose/activity | YOLO26 Nano Pose | 17 keypoints; conservative static arm-raised geometry; walking/running/waving remain unknown |
| Faces | MediaPipe Face Landmarker | Observable blendshapes and tentative smiling-looking result for sufficiently large faces; no identity or inner-emotion inference |
| Depth | Depth Anything V2 Small | Relative heatmap; no metric distance |
| Walkable candidates | SegFormer B0 / Cityscapes | Sidewalk mask with confidence threshold, uncertain road regions, detected-object boxes marked as obstructions; experimental, not navigation certification |
Face outputs are independent of person boxes and are not identity-linked. Small/occluded faces often produce no landmarks or unknown expression. No general-purpose temporal action classifier or SLAM trajectory estimator is implemented. Segmentation does not fuse depth into a calibrated traversability model; never drive a physical robot solely from these outputs. The default detector is optimized for speed; the full CPU pipeline is substantially slower.
Troubleshoot a local change
- Reproduce the smallest example from the quick-start guide.
- Compare required configuration and dependency versions with the README.
- Check the linked issue tracker for the same error. Include the command, runtime version, and relevant error when reporting a problem; omit credentials.
Source and help
The catalog identifies the license as AGPL-3.0. Read the repository license before redistributing source or assets.
Source captured: 2026-10-11
