SCM brings local AI search to Mac photos and video
A new MIT-licensed Mac app indexes photos and video frames with on-device CLIP models, so searches run offline after one 435 MB download.
In short
SCM is an open-source Mac app that makes photos and individual video frames searchable by meaning, using vision models that run entirely on the machine.
At a glance
- Four vision models: CLIP ViT-L/14@336 (default, about 435 MB), SigLIP-2-B/16, SigLIP-2-L/16@256 and SigLIP-B/16@384.
- Five video sampling presets, from 60 seconds down to 2.5 seconds per point; ffmpeg detects shot boundaries.
- Transcripts via Whisper tiny.en (about 150 MB) or base.en (about 300 MB); OCR via Tesseract with English plus 35 optional languages.
- Optional local chat through a llama.cpp sidecar running Qwen3 1.7B (about 1.1 GB) or Llama 3.2 3B (about 2 GB).
- Requirements: macOS 12 or later on Apple Silicon, Bun as the runtime; MIT license, version 0.2.4.
SCM, short for Screen Memories, answers the kind of question a filename index cannot: find the shot where someone holds a red umbrella. The MIT-licensed Mac app embeds photos and sampled video frames with vision models that run on the machine itself, then ranks results by meaning instead of by name.
Five ways into a library
File search ranks by visual meaning, with filename and phrase boosts layered on top. Scene search cuts a video into shots using ffmpeg boundaries and hands back timecodes rather than whole files. OCR covers text baked into images and frames, dialogue search matches exact words in Whisper transcripts, and an optional local chat sits on top of all three.
The model stack
Vision runs through ONNX Runtime with four interchangeable encoders: CLIP ViT-L/14@336 by default at about 435 MB, plus SigLIP-2-B/16, SigLIP-2-L/16@256 and SigLIP-B/16@384. Speech uses Whisper tiny.en (about 150 MB) or base.en (about 300 MB); OCR uses Tesseract, with English always enabled and 35 further languages available as 2.4 to 5 MB packs. The chat sidecar is llama.cpp, loading Qwen3 1.7B (about 1.1 GB) or Llama 3.2 3B (about 2 GB), and the project says weights download once and everything afterwards stays offline.
Indexing density is the dial
Sampling presets run from Eco at one point per 60 seconds to Ultra Pro at one point per 2.5 seconds, with Balanced at 30 seconds, Detailed at 15 and Ultra at 5 in between. Finer sampling catches shorter moments and costs index time and disk space. Everything the app derives, including index JSON, Float32 embedding bins per model, scene and transcript sidecars, thumbnails and posters, lands under ~/Library/Application Support/scm.
What is not established yet
The repository is candid about rough edges: embedding versions cap at ten snapshots, files that fail are retried three times and then skipped until they change, and unsigned builds run locally but can trip Gatekeeper. Requirements are macOS 12 or later on Apple Silicon with Bun as the runtime, installed through a Homebrew cask. No independent benchmark exists for indexing throughput or retrieval quality, so every figure here comes from the project's own documentation for version 0.2.4 in the repository, not from a test we ran. The Show HN entry stood at 141 points and 66 comments in our feed snapshot; we did not read the discussion thread itself, and the repository showed 289 stars and 14 forks when we fetched it.
FAQ
Does SCM upload my photos to a cloud service?
No. The project states that model weights download once and everything after that runs offline, with no account, no uploads and no telemetry.
What do I need to run it?
An Apple Silicon Mac on macOS 12 or later, plus Bun. The default vision model adds about 435 MB; the optional chat models add up to 2 GB.
Can it search what people say in a video?
Yes. Local Whisper transcripts (tiny.en or base.en) back an exact-word dialogue search, while scene search returns timecodes for the matching shot.