LensVLM-9B by Apple

LensVLM-9B by Apple

Apple has unveiled LensVLM‑9B, a 9‑billion‑parameter vision‑language model that can generate captions, answer visual questions, and perform other multimodal tasks. The model is built on a transformer architecture and trained on a large image‑text dataset. Apple claims it can be integrated into its ecosystem for improved photo search and accessibility features. The release was announced on Hacker News.