Python API
from ai_vision import VisionConfig, VisionEngine
vision = VisionEngine(VisionConfig(ocr=True, pose=True))
result = vision.analyze_image("street.jpg")
print(result.model_dump_json(indent=2))
# For a real video stream: BGR uint8 frame, monotonic media timestamp in seconds.
result = vision.analyze_frame(frame, timestamp_seconds=1.2, frame_id=36)
vision.reset_tracking() # before switching to another streamThe API accepts filenames or BGR NumPy arrays. Pydantic models validate serialized results. Unknown scene fields are null when segmentation is disabled; static movement is null. Heatmaps live in vision.artifacts and the result records the model identifiers and enabled/disabled modules.
Each analyzed image gets original, object, OCR, pose, combined JPEGs and JSON; enabled depth and segmentation add separate images. Missing module images are deliberately absent when disabled. OCR/pose-only views can equal the original when no prediction is available. The film animates the independent photographs and reveals model overlays; its zoom is presentation motion, not inferred scene motion.
Source captured: 2026-10-11