Finally trying out Agentic stuff and now the question is how do I get my text in the real world in to my computer… OCR, Vision, etc.

Wait 5 sec.

I’ve come to realize I’ve captured a lot of text with my phone. Sure it has auto OCR on phone but by the time it’s on my computer it’s gone. I rather automate the process. Debating between vision and OCR or both with agentic harness voting on which is likely better. I currently plan to use pi.dev on a folder in my computer to convert each image into text then stitch it into one final document. I’d also like some sort of image model that can take a crappy cell phone image and turn it into a nice document that looks scanned. It seems image edit models can work to some degree. What are you all doing? What tools / models do you recommend?   submitted by   /u/silenceimpaired [link]   [comments]