LensVLM: Compressing long context as images, expanding only relevant pages
LensVLM is a new approach that turns lengthy textual contexts into compact image representations, allowing models to store large amounts of information efficiently. When a user queries the system, only the image segments that are relevant to the question are expanded back into text, reducing computational load while preserving accuracy. This technique promises faster, more scalable language‑model interactions.