I have been building a tool that checks whether a message is a scam. Before letting anyone near it I put it through about 1,800 test cases: text message scams, email scams, prompt injection attempts, a holdout set the tuning never sees. Last week a real one got through, and it was the dumbest possible attack. The whole email was two words, "thank you", and one 416 KB JPEG. The image was a fake Geek Squad renewal notice, $369.99, with a phone number to call if you wanted to cancel. No link to scan, no attachment to sandbox, no text to read. My detector scored it LOW. The actual email, redacted: https://sentinelvault.net/static/images/tested-1800-scams-a-picture-got-through.jpg Here is the actual root cause, and it is the part I think generalizes. Every one of those 1,800 test cases was typed-out text. Not one had an image, not one had a PDF, not one had a real HTML email body. So when I changed how inline images were handled a few weeks earlier, the change was invisible to every gate I had. All 1,800 still passed, because none of them could tell the difference. The test suite was only ever testing what I had imagined. Two things came out of fixing it: Inline images that ARE the message now get read, instead of being treated as page furniture like a signature logo. More useful: when it cannot read something, it now says so and floors the verdict at medium instead of low. It used to tell you a scam was low risk. Now it tells you nobody looked. One other thing worth sharing, because it surprised me. The fix was to have a model read the image and pull out the text. I tested it against images with no text in them, because real mail is full of logos and family photos. Asked to return nothing for a text-free image, the model described it instead. "I can see this image appears to be mostly noise." Four times out of four. That description then became, as far as the rest of the system was concerned, the words in the image, and on one genuine email it was close enough to suspicious that it tripped my prompt-injection check. I was flagging a real person's photo because my own model narrated it. Rewriting the reader to return a structured yes-or-no plus the text dropped that false warning from three of my stored real messages to one. Full write-up with the numbers, including what survives when someone forwards you a suspicious email (spoiler: two thirds of forwards arrive forensically untraceable): https://sentinelvault.net/insights/tested-1800-scams-a-picture-got-through?utm_source=reddit&utm_medium=social&utm_campaign=picture_got_through_2026_10&utm_content=sideproject The thing itself is Scam Sentinel: https://scamsentinel.app/?utm_source=reddit&utm_medium=social&utm_campaign=picture_got_through_2026_10&utm_content=sideproject_app Being straight with you about where it is, since this sub has seen enough launch pages. It is pre-launch. iOS is on TestFlight, Android is close behind, and the site is an early access list rather than a download right now. So that link is a signup, not a product you can go use tonight. The write-up above is the part with actual substance in it. Happy to answer anything about the detection design. The part I am least sure about is where to draw the line on "cannot read it" so it fails safe without becoming noisy.   submitted by   /u/WestCoast_Pete [link]   [comments]