Sam Altman reportedly told OpenAI interns that within six months, a descendant of ChatGPT could watch a user's screen, record meetings and calls, and hold near-complete context about their life.
The statement is a forecast from an interview, not a product announcement or committed launch date. The timeline is aggressive. The underlying pieces are less distant than they sound.
Most components already exist
Screen sharing, meeting transcription, multimodal understanding, long context, memory, tool use, and background agents already work in bounded products. Combining them does not require one scientific breakthrough.
The gap is integration. A useful system has to follow activity across apps, identify what matters, connect conversations with files and decisions, retrieve the right history, and avoid confusing incidental screen content with durable truth.
Reliability becomes harder as context expands
More context can reduce repeated explanation. It can also add stale, private, contradictory, or irrelevant information. A model that sees everything still needs to know what it may use, what it should forget, and when it must ask.
Meeting capture adds another problem: the user is not the only person in the data. Consent, speaker identity, retention, access, and organizational policy have to work before the memory feels dependable.
Six months is plausible for a bounded product
A constrained version could arrive quickly: opt-in sessions, supported apps, visible recording, clear memory controls, and limited actions. A trustworthy system with complete life context is a much larger claim.
Read the forecast as a direction, not a deadline. AI products are moving from isolated conversations toward continuous context. The capability is close enough to plan for, but not mature enough to treat as inevitable, complete, or safe by default.
