Unveiling the Value of Motion for
Cinematic Camera Trajectories

NeurIPS 2026
Ziqi Zhou1, Yujian Yuan2, Laura Sevilla-Lara1
1University of Edinburgh, 2The Hong Kong University of Science and Technology
CineScript construction and the Pose9D versus DirSpeed representations.

Movie clips are re-processed into camera trajectories, motion captions, screenplay-style loglines and linked movie attributes (left). Instead of per-frame poses (Pose9D), we represent a trajectory by the direction and speed of its frame-to-frame motion (DirSpeed, right).

Abstract

Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation.

To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed improves alignment and generation while preserving measurable correlations with movie-level attributes.

DirSpeed: a motion-centric representation

Human descriptions of camera work say how the camera moves (“gradually dollies in and pans right”), not where it sits in a coordinate frame. Pose-centric representations make a model recover that motion implicitly. DirSpeed makes it explicit: from consecutive poses we take the translational step Δtt and the relative rotation ωt, and split each into a unit direction and a log-speed,

xtds = [ dttr, dtrot, sttr, strot ] ∈ ℝ8.

Normalising isolates the direction of movement from its magnitude, and the logarithm compresses the large dynamic range of real camera speeds. Given the first pose, a DirSpeed sequence integrates back to per-frame poses deterministically.

Alignment with text

We evaluate alignment with a lightweight, purely contrastive protocol that drops the reconstruction objective of CLaTr [1]. Under both evaluators, switching the trajectory representation from Pose9D to DirSpeed improves retrieval, and in the joint embedding spaces matched trajectories (•) and captions (★) fall into shared clusters instead of staying apart.

CLaTr with DirSpeed

CLaTr · DirSpeed

CLaTr with Pose9D

CLaTr · Pose9D

Ours with DirSpeed

Ours · DirSpeed

Ours with Pose9D

Ours · Pose9D

EvaluatorRep.R@1 ↑R@5 ↑R@10 ↑MedR ↓#Params
OursDirSpeed25.240.749.0113.6 M
Pose9D17.826.531.4513.6 M
CLaTrDirSpeed19.730.539.22330 M
Pose9D6.912.014.736130 M

The CineScript dataset

CineScript contains about 28K movie clips and 10M frames (12.0 s per clip on average), drawn from ShotBench [2], CineTechBench [3], MovieShots [4], CMD [5] and the film subset of VADB [6] and re-processed through one pipeline. Each clip carries a ViPE [7] camera trajectory, a motion caption, and a screenplay-style logline ([INT./EXT.] [Location] – [Time] – [Action]). A subset of 3,163 clips from roughly 1,400 films is linked to real movie metadata: release year, genre and director.

Motion primitive distribution.

Distribution of translation (yellow) and rotation (green) primitives.

Dolly-in clips with different movie attributes.

The same coarse motion (a dolly-in) executed differently across movie attributes.

Probes trained only on real trajectories predict era, genre and director above chance, and do so more reliably from DirSpeed than from Pose9D. These are diagnostic probes of latent correlation, not a claim of style control.

Attribute#ClassDirSpeedPose9DΔ
Era349.847.3+2.6
Genre362.160.7+1.4
Director368.251.8+16.4

Macro-F1 (×100) on real validation trajectories.

CineGEN: masked autoregressive generation

CineGEN architecture and the contrastive evaluation protocol.

Diffusion generators such as CCD [8] and E.T. [1] refine every position in parallel, while GenDoP [9] decodes causally in a fixed order. CineGEN is a text-conditioned masked autoregressive model [10] that works directly on continuous DirSpeed tokens: a Transformer sequencer reads the partially masked trajectory, and a small diffusion head samples the missing tokens. Because trajectory features are already low-dimensional, no autoencoder is needed. The model is conditioned on the motion caption, the logline and the first camera pose, and reveals positions in order of contextual confidence, so it commits to the motion progressively while keeping bidirectional context over the whole shot.

Results

We compare with CCD [8], E.T. [1] and GenDoP [9], each both as the released checkpoint (∗) and retrained on CineScript in its native formulation. CineGEN is best on every metric; against the strongest retrained baseline it raises R@1 from 0.93 to 3.35 and reduces FCD by nearly 70%.

Method Trajectory quality Text–trajectory alignment Movie attributes
F1 ↑FCD ↓Cov. ↑ AlignScore ↑R@1 ↑MedR ↓ Era ↑Genre ↑Director ↑
CCD∗0.11752.450.2984.900.081032.544.346.312.1
CCD0.12849.190.27018.820.43397.047.256.332.7
E.T.∗0.015135.740.0340.000.08960.040.756.725.0
E.T.0.19416.390.59629.780.81251.042.057.528.1
GenDoP∗0.175102.970.0857.630.23715.533.545.618.1
GenDoP0.23422.100.58633.110.93195.033.252.529.9
CineGEN (Ours)0.4376.770.78357.793.3552.047.860.748.1

∗ released pretrained checkpoint; unmarked baselines are retrained on CineScript.

Qualitative comparison

Generated trajectories from CineGEN and the baselines for two captions.

CineGEN produces smoother, more coherent paths that follow compound instructions. Starred columns are the released checkpoints.

Re-rendering real footage

To see the trajectories as camera work, we re-render CameraBench [11] clips with CameraAnything [12], a camera-conditioned video generator. The scene, the renderer, its settings and the seed are fixed; only the trajectory, generated by each method from the same motion caption, changes. In a blinded study with 24 participants (720 judgements), CineGEN was selected as following the reference motion in 70.1% of judgements and was the only version selected in 30.1%, against 28.9% and 6.2% for the strongest baseline.

Selection and sole-selection rates of each method in the blinded human study.

Examples

Each row re-renders the same clip with CameraAnything [12]; only the camera trajectory differs. Videos play when scrolled into view.

BibTeX

@inproceedings{zhou2026unveiling,
  title     = {Unveiling the Value of Motion for Cinematic Camera Trajectories},
  author    = {Zhou, Ziqi and Yuan, Yujian and Sevilla-Lara, Laura},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}

References

  1. R. Courant et al. E.T. the Exceptional Trajectories: Text-to-camera-trajectory generation with character awareness. ECCV 2024. arXiv:2407.01516
  2. H. Liu et al. ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models. NeurIPS 2025. arXiv:2506.21356
  3. X. Wang et al. CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation. NeurIPS 2025. arXiv:2505.15145
  4. A. Rao et al. A Unified Framework for Shot Type Classification Based on Subject Centric Lens. ECCV 2020. arXiv:2008.03548
  5. M. Bain et al. Condensed Movies: Story Based Retrieval with Contextual Embeddings. ACCV 2020. arXiv:2005.04208
  6. Q. Qiao et al. VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations. NeurIPS 2025. arXiv:2510.25238
  7. J. Huang et al. ViPE: Video Pose Engine for 3D Geometric Perception. CVPR 2025. arXiv:2508.10934
  8. H. Jiang et al. Cinematographic Camera Diffusion Model. Eurographics 2024. arXiv:2402.16143
  9. M. Zhang et al. GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography. CVPR 2025. arXiv:2504.07083
  10. T. Li et al. Autoregressive Image Generation without Vector Quantization. NeurIPS 2024. arXiv:2406.11838
  11. Z. Lin et al. Towards Understanding Camera Motions in Any Video. NeurIPS 2025. arXiv:2504.15376
  12. Y. Li et al. CameraAnything: Refilming Videos with Arbitrary Camera Control. ECCV 2026. arXiv:2607.24591