AI can answer increasingly complex questions about the world, yet it still knows remarkably little about the life of the person asking them. It doesn’t inherently know what happened in your day, which moments mattered, or how an experience from last week connects to something happening now. Intelligence without the surrounding world falls short of its full potential.To help us in everyday life, AI needs to understand our experiences: what happened, what mattered, and how different moments relate to one another. Useful memory goes beyond recalling an event; it connects the dots.A common approach to replicating this via AI is to extract facts from users’ chat histories and retrieve them when relevant. This is useful, but chat conversations alone provide an incomplete picture of someone’s everyday life. People describe experiences selectively, speak hypothetically, and express opinions that may change. A system can mistakenly turn “I’m thinking about becoming vegetarian” into a lasting assumption: “This person is vegetarian.”So how can we construct memory systems that help AI understand our everyday experiences? At Looki, we’ve identified four critical challenges:1. Deciding what to rememberIf you think about how our own memory works, our eyes take in our surroundings, but we don’t “log” everything we see. You remember being in a grocery store, but not every item that you passed. Similarly, everyday audio and video contain noise, repetition, and seemingly unimportant details.Processing a continuous stream of images and audio requires substantially more computation—and consumes far more AI tokens—than processing chat text, while sustained daily use requires power-efficient hardware and software.At Looki, we compress memory at multiple levels of detail. For a commute, this would look like:Level 1 — Main activity: The user commuted from home to work.Level 2 — Events: The user walked to the subway station, took the subway, and walked from the station to the office.Level 3 — Event Details: How many stops they traveled, what the weather was like, and other details within those moments.Different questions require varying levels of detail. “What did I do this morning?” needs a broad overview; a question about one part of the journey requires more specificity. Temporary context, such as an afternoon plan, should become less influential as it becomes less relevant, while repeated observations can form a clearer understanding of routines or preferences. Long-term personal traits and facts, such as how many children someone has, should persist, while remaining correctable as circumstances change.We optimize precision and recall at each level: whether retained information is accurate and whether relevant information is preserved. A system might recognize a commute while missing an important detail, so evaluating only the high-level summary would hide that failure.2. Identifying the right memoryThe next challenge is identifying the right information when a memory is “recalled.” People remember fragments, not exact timestamps.To function like memory, AI must piece together clues from conversations, places, and times as the collection of memories grows. When the evidence is missing, it should acknowledge the gap.Take the question, “What was the restaurant someone recommended after our meeting?” The system needs to take the following steps:Interpret the question. Identify the clues: a meeting, subsequent conversation, and restaurant recommendation.Find candidate moments. Looki combines semantic search with keyword search, then reranks candidates by relevance.Expand the search. Examine moments that are close in time and semantically similar.Resolve the reference. Determine which restaurant was recommended, distinguishing a recommendation from a passing mention.Check the evidence. Load the original visual and audio data from relevant moments to ground the answer.Answer or clarify. Return the answer with supporting context, or ask a focused question if the evidence still leaves multiple possibilities.Failure can happen at several stages:Capture: A camera may face away from the relevant object, or background noise may obscure a name; reasoning cannot recover what was never captured.Compression: A summary may preserve a moment but omit a detail the user needs later.Retrieval: Similar routines can produce similar moments, causing the system to retrieve the wrong one.Interpretation: Even with the right moment, the LLM might confuse a suggestion with a decision, or a mention with a recommendation.Checking original visual and audio evidence can resolve those distinctions, but at the cost of additional computation and response time. The challenge is deciding how much investigation a question warrants, while recognizing when the evidence is insufficient.3. Connecting the dotsConnecting experiences over time could help users notice recurring concerns, revisit unfinished ideas, or understand how their interests have changed. This lays the groundwork for what Tulving called episodic memory: memory about concrete personal experiences, such as “what did I eat before my flight?” or “who was I sitting with during that meeting?” Semantic memory, conversely, stores general facts like“Paris is the capital of France.”How AI surfaces these connections depends on the need: a short answer, original clip, written reflection, or video recap. Any insight should explain which experiences support it and distinguish between captured evidence, user statements, and AI interpretations. An AI-generated conclusion should not become more trustworthy simply because it appears again in a later summary.4. Learning from corrections and changesMemories can contain mistakes, and people change. A user may correct someone’s name or explain that an old preference no longer applies. For AI to be truly useful in everyday life, it must carry those updates into future answers while distinguishing between “this was wrong” and “this used to be true.”Our approach proactively detects conflicts between existing memories and new information, then asks the user to resolve them through chat or by editing information in the app. For example, several visually captured meals might lead the system to infer that the user is vegetarian. If later evidence conflicts with that assumption, it can ask the user to clarify, rather than silently choosing one interpretation. That feedback corrects the broader assumption without changing what was observed in individual meals.Storing a new statement is relatively straightforward; deciding whether it corrects, replaces, or adds context to an older memory is harder. Without a reliable feedback loop, outdated assumptions can keep appearing in future answers.A basic RAG setup retrieves relevant stored passages, but relevance alone does not determine whether a statement is still valid. Managing time, contradictions, and dependencies requires recording a correction, updating affected summaries and indexes, and preventing invalidated claims from returning through older cached answers. RAG can support this, but these behaviors require explicit design.We’re building memory primarily around visual evidence from everyday experiences, supported by audio, chat, location, and other context. This gives memories a basis in what we observed. Because that evidence can still be incomplete or misinterpreted, ongoing recordings can confirm or revise earlier information, while proactive AI can ask users for clarification. Together, these inputs allow the system’s understanding to evolve as it learns more and the user’s life changes.The promise of personal AI is to help us make better use of our experiences. That requires memory we can understand, check, and correct—and insights that remain grounded in what actually happened.