Deviation–Consequence Preview
Pairs a concise “What if…?” cue with a predicted adverse outcome so users can understand a mistake before making it.
Generative Interfaces · Augmented Reality
The University of Hong Kong · University of California, Los Angeles
2026
Procedural guidance can show users how to act without explaining what an action enables later. We present Foresee, an interface combining next-step guidance with consequence previews, method comparisons, and cross-step access. A hybrid pipeline combines structured rendering, image editing, and video generation. A literature-informed framework and six-participant formative study shaped its design. A 24-participant study compared AR guidance, AR with generated video, and Foresee across digital painting, gameplay, and cocktail making. Exploratory painting comparisons showed higher correct completion without a recorded error than AR and higher post-task quiz accuracy than generated-video guidance under task-wise correction. Interviews revealed that users could overlook valued explanations, match visible outcomes without understanding later possibilities, or wait for concrete demonstrations when symbolic cues were unclear. These findings inform generative previews that explain what actions enable or constrain while supporting ongoing execution.
Foresee integrates three complementary views of the future into ordinary next-step guidance.
Pairs a concise “What if…?” cue with a predicted adverse outcome so users can understand a mistake before making it.
Compares feasible methods through their outcomes, costs, affected steps, and action demonstrations.
Provides a compact map of task progress with access to generated and upcoming previews without changing the active step.
The AR client supplies the task request and current scene. Foresee grounds the task, generates structured UI and scene-conditioned media, and returns spatial cues, action videos, and outcome images.
The hybrid pipeline uses structured rendering for rule-defined content and diffusion models for complex appearance and motion. Reviewed procedural plans separate each action from its resulting state, enabling the system to generate both instructions for how to act and previews of what the action may cause.
A 24-participant study compared Foresee with AR guidance and AR with generated video across digital painting, gameplay, and cocktail making. The evaluation examined task execution, post-task procedural understanding, usability, workload, and how participants chose to inspect future previews.
In exploratory digital-painting comparisons, Foresee produced higher post-task quiz accuracy than generated-video guidance and more correct steps without a recorded error than AR.
@article{sun2026foresee,
title = {Foresee: Making Procedural Futures Tangible through Generative Previews},
author = {Sun, Yusi and Jiang, Ying and Lu, Jiayin and Qian, Jing and Yang, Yin and Kuo, Yong-Hong and Jiang, Chenfanfu},
year = {2026}
}