Turn feelings into visuals: understand photos, discover potential, and find the right visual language.
#computer-vision
13 posts
[dsh]为纯文本模型设计更强大的视觉工具箱:安装免费使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integrat...
A Skill and CLI tool for Codex and Claude Code that converts images, PDFs, image-based PPTX files, and mixed PPTX files into editable PowerP...
Agent skills for Ultralytics YOLO, covering models, datasets, training, tuning, inference, tracking, and export.
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents...
A unified framework for tabular, time-series, and multimodal machine learning
Local TradingView chart analysis via OpenCV + Tesseract OCR. No API keys, no GPU, runs fully offline — MCP server with color/level/zone dete...
OCR in reverse. Make your documents worse. Python toolkit for synthetic logistics document & photo datasets with verified ground truth.
Zero-token, on-device vision for AI agents — OCR, pixel-exact UI targeting, and face grouping on Apple Vision. Pure Swift, single binary, no...
A multi-interface (REST and MCP) server for automatic license plate recognition 🚗
MCP for Ultralytics Platform workflows, datasets, training, prediction, and model operations.
Open-source Hardware AI agent. Single Rust binary for cameras, sensors, robots, and IoT fleets — orchestrated by AI agents with memory and r...