VLM

In Preparation: CVPR 2027

Hawkeye

Brian Tang, Kang G. Shin

A vision-language model fine-tuned to read and reason about illegible text degraded by motion blur, low resolution, occlusion, and glare. We introduce BSTRB, a real-world blurry scene-text benchmark, and Hawkeye4B, a 4B-parameter model that rivals much larger proprietary VLMs at reading noisy in-the-wild text. Contribution: lead author.

blog

// connect

Follow or contact me

I publish and open-source my work. I also occasionally post random thoughts.