Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, potentially democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user’s appearance. On the other hand, self-reenactment methods can render coherent 3D portraits for telepresence applications by driving a personalized 3D prior, but fail to faithfully reconstruct the user’s per-frame appearance (e.g. facial expressions and lighting). In this work, we recognize the need to maintain both personalized stable appearance and dynamic video conditions to enable the best possible user experience. To this end, we propose a new fusion-based 3D portrait reconstruction method, which captures the authentic dynamic appearance of the user while fusing it with a personalized 3D subject prior, producing temporally stable 3D videos with consistent personalized appearance and structure. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction accuracy and temporal consistency on in-studio and in-the-wild datasets.