Hugging Face's new iOS app can identify and describe whatever is in your camera's field of view
Hugging Face has released HuggingSnap, an iOS app that runs SmolVLM2 locally and efficiently to enable visual understanding and text-based generative AI tasks, such as answering questions about an image or identifying and describing objects from visual input.
Hugging Face recently launched HuggingSnap, an iOS application that runs SmolVLM2, a small but performant multimodal language model that accepts video, images, and text as inputs, and generates text in response. It can be used for:
- vision understanding tasks, such as answering questions about or identifying and describing objects within an image or video;
- generating text grounded on visual information, like writing a story grounded on the contents of one or more images;
- text-only tasks, as one would with a standard language model.
Apps that leverage multimodal generative AI for various vision and text-based tasks are hardly new. However, HuggingSnap's key selling point is that SmolVLM2 is run locally and efficiently. As a result, the app does not require an internet connection to work, and all data is processed within the device without performance losses.
HuggingSnap can be downloaded from the App Store or built from the GitHub repository. It requires an iPhone running iOS 18 to run.
Ellie Ramirez-Camara is the News Editor at Data Phoenix, where she writes the daily AI newsdesk — covering model releases, research, funding rounds, and policy across the AI and machine-learning industry. She tracks announcements from labs and startups alike and distills them into clear, source-linked reporting for practitioners.
More news

xAI releases Grok Voice Transcribe 2.0 for its speech-to-text API

OpenAI proposes six-pillar youth safety blueprint for Australia

Google confirms Gemini accessed three companies during a safety test
