came across a pretty interesting technical breakdown of chatgpt's newly launched "intelligent ui" feature, and thought this subreddit might find it interesting. for anyone unfamiliar with the concept, intelligent ui is essentially openai's take on generative ui. instead of restricting llm responses to plain text or markdown, the model can compose actual interactive interfaces in real time. there are different approaches to making this work. some systems let the model choose and compose elements from a predefined component library, while others allow it to generate entire interfaces on the fly (basically writing html/react code and rendering it inside an iframe). it's more of a spectrum than a single technique. projects like openui, vercel's json-render, google's a2ui, and now chatgpt's intelligent ui all sit somewhere along this spectrum, with different trade-offs in flexibility, reliability, performance, and how much freedom the model gets. but that's not even the most interesting part. These folks managed to reverse engineer chatgpt's implementation in less than 24 hours after launch! what's particularly impressive is that they claim to have done this entirely through publicly observable behavior, without access to openai's internal codebase. from their write-up: “All observations come from our own ChatGPT accounts, from the traffic the ChatGPT web app generates, and from the JavaScript that chatgpt.com serves publicly.” found this pretty fascinating from an engineering perspective, especially considering how quickly they managed to put together a breakdown of how the system works. and then there's the funnier part. the same team released something called open intelligent ui, which is a pretty obvious jab at how openai isn't really "open" anymore. the joke works even better when you realize these guys actually own the domain openui.com lol. the idea they're pitching is that you can recreate experiences similar to chatgpt's new intelligent ui inside your own applications using their open source framework. and here's where it gets particularly interesting, you can technically do all of this with local llms. since openui is model agnostic, you can integrate it with local models through ollama, lm studio etc. it's not necessarily a one click, out of the box recreation of chatgpt's experience, but from what i understand, the underlying pieces are there to build something similar that runs entirely locally. i initially came across these folks through a viral twitter post comparing chatgpt's intelligent ui with openui's generative ui, and ended up going down a rabbit hole reading about the different approaches to generative ui. some helpful links for anyone interested: openui: https://github.com/thesysdev/openui open intelligent ui: https://github.com/thesysdev/open-intelligent-ui decoding chatgpt's intelligent ui: https://x.com/rabi_guha/status/2108238432355123572 chatgpt intelligent ui vs openui generative ui: https://x.com/KashyapVisharad/status/2108239102676250933 json-render: https://github.com/vercel-labs/json-render a2ui: https://github.com/a2ui-project/a2ui would love to know what everyone here thinks about generative ui in general. is this actually a useful direction for llm interfaces or is it another one of those things that looks amazing in demos but doesn't translate particularly well to real world applications? i'm especially curious about the local inference angle. with smaller models getting increasingly capable, do you see a future where something like this becomes practical entirely on device? or is the additional complexity, latency and structured output overhead simply not worth it compared to a conventional ui? local llama has been my go to subreddit for years whenever i come across something interesting in the llm space, so genuinely curious what the general opinion here is. would love to hear your thoughts, especially if you've tried building something similar with local models!   submitted by   /u/Mr_BETADINE [link]   [comments]