NextMe-800: Anticipating Personal Behavior
from Months of Egocentric Video

If a system watched four months of your life through your own eyes, could it predict what you'll do next?

Explore the moments
Hover to step closer · click to hold a moment

01 / The question

Plans meet real life.

We make ambitious plans all the time: learn the piano, start a workout routine, finish that research project. Actual behavior rarely follows the script. We get distracted, change priorities, and fall back into familiar routines.

Hindsight is remarkably useful. Looking back, we can often see where our plans diverged from what we actually did. But hindsight arrives after the fact. What if a system could preview plausible paths before you commit, based on your own history?

That is the question behind NextMe-800: not what a generic person might do, but what this person might do next, given the patterns of their life.

02 / One life, observed over time

795.5 hours.
126 calendar days.

To study personal behavior, we need temporal depth within one person. NextMe-800 records 1 Hz first-person RGB, gaze, and audio across four months, from research and meals to travel and rare encounters.

528recording sessions
39.5%estimated awake-time coverage
18days with no recording
Daily recorded hours over 126 calendar days, including blank days and a seven-day moving average
Figure 2a · Daily recorded hours. Breaks for charging and incompatible activities remain real gaps in the record. Open full-size figure

Ego4D includes 931 participants but averages about 3.9 recorded hours per person. EgoLife follows six people for seven days. NextMe-800 complements that breadth with months of dense observation of one person.

Its scenes range from coding, study, meals, exercise, dating, and social time to dolphins, karst landscapes, and robotic dogs. This is a curated view of habitual activities and uncommon events, not an unbiased sample of their frequency.

03 / From pixels to behavior

Zoom out without losing the moment.

A three-hour study session contains thousands of atomic actions: editing code, checking a phone, adjusting the air conditioner. Thirty-second video windows become first-person captions; a five-level hierarchy then groups them by shared goals.

Paper Figure 3 showing one puzzle session at low, mid, and high levels of the caption tree
Figure 3 · One puzzle session at three of five abstraction levels. Open full-size figure
L1529,431 events~5 sec
L248,093~1 min
L36,367~7.5 min
L41,988~24 min
L51,236~39 min

428× fewer events from L1 to L5. Each step removes detail while preserving a broader goal.

One recording, three resolutions.

A 3 h 37 min recording, from major activity to atomic detail.

L5 2 eventsL2 178 eventsL1 2,287 events
3 h 31 min

I work on RL homework

… earlier tasks …

I compute target Q-values

I edit target_q = …

…

I scratch my nose

I glance at my phone

……

I chat on WeChat

I open a WeChat notification …

I browse video recommendations

I view video recommendation cards …

I adjust the AC

I activate AC display

I increase temp. to 26°C

… later tasks …

Table 6 · One major activity contains many tasks; each task contains finer actions.

L1 keeps the nose-scratch and phone glance. L5 turns thousands of details into a few major activities. Neither scale is intrinsically best: the right level depends on whether the forecast concerns a moment or the rest of an afternoon.

At each level, the hierarchy abstracts adjacent events sharing a goal. It does not simply cut the day into fixed-duration blocks.

04 / Measuring an open-ended future

What counts as a match?

Suppose the recorded action is “I grab a bottle of cola from the fridge,” while a prediction says “I pick up a glass of water from the table.” Exact word overlap misses some meaning; a loose semantic score may overlook the changed action and object.

01 · A single action

Embedding similarity

Qwen3-Embedding-8B compares the meaning of two atomic actions, including their core verbs and objects.

02 · A sequence

Semantic edit distance

Soft edit distance (SED) aligns whole trajectories. A substitution costs 1 − embedding cosine similarity; insertions and deletions cost 1. The sequence score is normalized against chance.

Human agreement and score spread from Table 2
MethodPairwise ↑Spread H₁₀ ↑
Semantic models
Qwen3-Reranker-8B0.94170.6429
Qwen3-Embedding-8B0.92900.9724
Language-model judges
Gemini-3.7-Flash0.90970.9407
GPT-5.60.89680.9113
Claude-Haiku-4.50.89520.9612
Qwen3.5-Flash0.87940.8852
Classic text metrics
BERTScore0.82690.5435
ROUGE-L0.76980.8037
BLEU0.69160.2878

The reranker agrees slightly more often with people but clusters many scores at the extremes. Embeddings preserve graded differences for SED. H₁₀ measures score spread, not correctness; these are semantic scores, not classification accuracy.

05 / How much history?

More memories don't always help.

Can in-context learning keep scaling when it sees more of one person's history? We vary the context from 1 to 3,000 recent events and ask how the model's semantic prediction score changes. Fine-grained scores rise quickly, then plateau or decline; coarser activities benefit from longer context, but eventually saturate too.

Paper Figure 5: five context scaling curves on a logarithmic time axis from seconds to days
Figure 5 · History duration varies with abstraction level. Shaded bands show variation across three generations; diamonds estimate performance at 50 history events. Open full-size figure

Fine atomic actions often depend most on the last few minutes. Broader activities can draw on patterns that unfold over days. The graph suggests different useful horizons at different levels, not a universal context length.

A 50-event window lies in or near a strong range for the paper's in-context experiments, so it is used in the main benchmark. It is a practical default, not proof that 50 is optimal for every situation.

06 / The benchmark

Can a model anticipate you?

NextAct combines 1,000 points from NextMe-800 with 500 from six EgoLife participants. Given 50 recent events, learned predictors generate three candidate trajectories for the next one or ten recorded events. The reported score selects the best match to the ground truth, a generous best-of-three evaluation.

0.3431best reported K=1 score
Gemini 3.1 Pro
Selected Table 3 predictors · normalized embedding scores across L1–L5
PredictorL1L2L3L4L5Mean

Frontier models cluster around 0.32–0.34 for the next event, above the repeat-recent baseline of 0.1916. For ten events, the best score is 0.1943. These are normalized semantic sequence scores, not percentages of exactly correct actions.

Codex + DeepSeek can retrieve history beyond the 50-event window, yet does not lead the reported mean at either horizon. Access to more past activity is different from knowing which past activity matters now.

A continuation that works

Recorded: “I read the Feishu document notes.”
Predicted: “I read the Feishu reflection notes.”

0.92 · K=1, L1
A transition it misses

Recorded: “I walk away from the table.”
Predicted: “I write mathematical formulas on the notepad.”

0.00 · K=1, L1

Examples come from the preprint appendix. The benchmark gives target recording windows; events during unobserved gaps are not treated as known behavior.

07 / The most useful scale

The middle is easier to foresee.

Atomic actions are noisy. Major activities are farther apart. Between those extremes, the hierarchy finds a more predictable scale.

Paper Figure 6: average semantic score rises from L1 to intermediate levels then falls toward L5, especially for ten-step prediction
Figure 6 · Frontier-model average by abstraction level; bands show variation across models. Open full-size figure

At K=1, the reported average rises toward L3–L4 (roughly 7.5–24 minutes per event) before a small decline at L5. The hierarchy removes incidental detail without yet losing all continuity.

At K=10, the peak is more clearly L3. Ten L5 events span 5.8 hours on average: by then, even an accurate sense of the person's routine may not identify the turns their day will take.

What comes next

So why does any of this matter?

Imagine a future where everyone has a digital twin. Before each day begins, we could explore a tree of possible futures — multiple trajectories of how we might actually behave — and choose the path we most want to live. By turning hindsight into foresight, personal behavioral prediction could help us make better choices, reduce regret, and face the future with greater clarity.

A forecast beside a recorded day

What would a preview change?

Recorded pathEarlier previewOther possibilities

Open a choice. Drag to explore. Zoom out to see the whole day.

100%

NextMe-800