When I started building Solofy, I expected the AI part to be the hardest. Choosing models, writing prompts, handling structured output, adding search, dealing with latency. That work definitely took time.But once the product started becoming real, completely different problems became harder. What happens when an API takes 30 seconds? What happens when the model returns valid JSON 99 times and breaks on the 100th? How do you keep AI costs under control when one feature can trigger multiple model or search calls? How do you make the UI feel fast when the backend is waiting on an LLM? And probably the biggest one: how do you make the AI output consistently useful instead of just technically correct? Building the AI feature was only one part. Making it reliable enough that someone can actually pay for it has been a much bigger engineering problem. I think this is something AI demos hide really well.   submitted by   /u/Obvious-Vast-1248 [link]   [comments]