View complete timeline
February 2019Model Release

Xoromancy

Aman Tiwari & Gray Crawford

Two Carnegie Mellon students built the first interactive, real-time tool that let the human body steer a neural network's imagination. Xoromancy — a fusion of the high-dimensional spaces inside machine learning models with "-mancy," divination — mapped the position, rotation, and fine motor movement of a visitor's hands, tracked by a Leap Motion sensor, directly into the latent space of BigGAN, DeepMind's state-of-the-art image generator of late 2018. Move your hands and the projected, pseudo-photographic image warps with you: textures, colors, and subject matter blending in real time. Gray Crawford announced it on February 28, 2019: "Traversing bigGAN's high-dimensional space of pseudo-real images is enthralling, like divination via movement." The design insight was that BigGAN's thousands of input variables were notoriously counterintuitive to navigate with sliders — but "the human body is itself a continuous and high-dimensional control system." Tiwari, the primary software engineer, built the pipeline in Unity, Python, and TensorFlow, surfacing fourteen components of the z-vector, seven per hand. As the project documentation put it: "Correlation between proprioception and visual feedback builds an intuitive understanding for navigating the highly-nonlinear mappings between input dimensions and generated output imagery." In other words: you learn to feel your way around latent space. Built in a month, it premiered at CMU's student-run Frame Gallery February 22–24, 2019, reached a public audience at New York Live Arts' Live Ideas festival that May, became a formal paper — "Xoromancy: Image Creation via Gestural Control of High-Dimensional Spaces" — at the IEEE GEM conference at Yale in June, and closed the year at INTERSECTIONS, the Frank-Ratchye STUDIO for Creative Inquiry's 30th-anniversary exhibition. Years before anyone typed a prompt, Xoromancy posed the question the whole field would inherit: not whether a model can generate an image, but how a human directs it — and answered, for one enthralling installation, with dance.

Made with BigGAN (DeepMind) + Leap Motion hand tracking