Paper reviews and research notes.
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow (CVPR 2026)
A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer (ICLR 2026)
Perfect linear concept erasure in closed form
Mask Diffusion Model for Scene Text Recognition (AAAI 2026 Oral)
Are VLMs Seeing or Just Saying?
Scaling Context Windows via Visual-Text Compression
Unlocking Test-Time Training in Vision
Neural Optical Understanding for Academic Documents
OCR-free Document Understanding Transformer
Feature Learning by Inpainting