Derek Nowrouzezahrai

Toshiya Hachisuka

Ravi Ramamoorthi

Tzu-Mao Li

Differentiable rendering is frequently used in gradient descent-based inverse rendering pipelines to solve for scene parameters – such as … (see more)reflectance or lighting properties – from target image inputs. Efficient computation of accurate, low variance gradients is critical for rapid convergence. While many methods employ variance reduction strategies, they operate independently on each gradient descent iteration, requiring large sample counts and computation. Gradients may however vary slowly between iterations, leading to unexplored potential benefits when reusing sample information to exploit this coherence. We develop an algorithm to reuse Monte Carlo gradient samples between gradient iterations, motivated by reservoir-based temporal importance resampling in forward rendering. Direct application of this method is not feasible, as we are computing many derivative estimates (i.e., one per optimization parameter) instead of a single pixel intensity estimate; moreover, each of these gradient estimates can affect multiple pixels, and gradients can take on negative values. We address these challenges by reformulating differential rendering integrals in parameter space, developing a new resampling estimator that treats negative functions, and combining these ideas into a reuse algorithm for inverse texture optimization. We significantly reduce gradient error compared to baselines, and demonstrate faster inverse rendering convergence in settings involving complex direct lighting and material textures.

2023-07-23

Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Proceedings (published)

Cone-Traced Supersampling for Signed Distance Field Rendering

Andrei Chubarau

Yangyang Zhao

Ruby Rao

Paul Kry

While Signed Distance Fields (SDFs) in theory offer infinite level of detail, they are typically rendered using the sphere tracing algorithm… (see more) at finite resolutions, which causes the common rasterized image synthesis problem of aliasing. Most existing optimized antialiasing solutions rely on polygon mesh representations; SDF-based geometry can only be directly antialiased with the computationally expensive supersampling or with post-processing filters that often lead to undesirable blurriness and ghosting. In this work, we present cone-traced supersampling (CTSS), an efficient and robust spatial antialiasing solution that naturally complements the sphere tracing algorithm, does not require casting additional rays per pixel or offline pre-filtering, and can be easily implemented in existing real-time SDF renderers. CTSS performs supersampling along the traced ray near surfaces with partial visibility identified by evaluating cone intersections within a pixel's view frustum. We further devise a specialized sampling strategy to minimize the number of shading computations and aggregate the collected samples based on their correlated visibility. Depending on configuration, CTSS incurs roughly 15-30% added computational cost and significantly outperforms conventional supersampling approaches while offering comparative antialiasing and visual image quality for most geometric edges.

2023-05-02

graphicsinterface.org/Graphics_Interface/2023/Conference_SD (published)

Visual Question Answering From Another Perspective: CLEVR Mental Rotation Tests

2023-04-01

Pattern Recognition (published)

MeshDiffusion: Score-based Generative 3D Mesh Modeling

Zhen Liu

Yao Feng

Michael J. Black

Liam Paull

Weiyang Liu

We consider the task of generating realistic 3D shapes, which is useful for a variety of applications such as automatic scene generation and… (see more) physical simulation. Compared to other 3D representations like voxels and point clouds, meshes are more desirable in practice, because (1) they enable easy and arbitrary manipulation of shapes for relighting and simulation, and (2) they can fully leverage the power of modern graphics pipelines which are mostly optimized for meshes. Previous scalable methods for generating meshes typically rely on sub-optimal post-processing, and they tend to produce overly-smooth or noisy surfaces without fine-grained geometric details. To overcome these shortcomings, we take advantage of the graph structure of meshes and use a simple yet very effective generative modeling method to generate 3D meshes. Specifically, we represent meshes with deformable tetrahedral grids, and then train a diffusion model on this direct parameterization. We demonstrate the effectiveness of our model on multiple generative tasks.

2023-02-01

ICLR.cc/2023/Conference (notable)

From Words to Blocks: Building Objects by Grounding Language Models with Reinforcement Learning

Michael Ahn

Anthony Brohan

Noah Brown

liang Dai

Dan Su

Holy Lovenia Ziwei Bryan Wilie

Tiezheng Yu

Willy Chung

Quyet V. Do

Paul Barde

Tristan Karch

C. Bonial

Mitchell Abrams

David R. Traum

Hyung Won

Le Hou

Shayne Longpre

Yi Zoph

William Tay … (see 32 more)

Eric Fedus

Xuezhi Li

Lasse Espeholt

Hubert Soyer

Remi Munos

Karen Si-801

Vlad Mnih

Tom Ward

Yotam Doron

Wenlong Huang

Pieter Abbeel

Deepak Pathak

Julia Kiseleva

Ziming Li

Mohammad Aliannejadi

Shrestha Mohanty

Maartje Ter Hoeve

Mikhail Burtsev

Alexey Skrynnik

Artem Zholus

A. Panov

Kavya Srinet

A. Szlam

Yuxuan Sun

Katja Hofmann

Marc-Alexandre Côté

Ahmed Hamid Awadallah

Linar Abdrazakov

Igor Churin

Putra Manggala

Kata Naszádi

Michiel Van Der Meer

Leveraging pre-trained language models to gen-001 erate action plans for embodied agents is an 002 emerging research direction. However, exe… (see more)-003 cuting instructions in real or simulated envi-004 ronments necessitates verifying the feasibility 005 of actions and their relevance in achieving a 006 goal. We introduce a novel method that in-007 tegrates a language model and reinforcement 008 learning for constructing objects in a Minecraft-009 like environment, based on natural language 010 instructions. Our method generates a set of 011 consistently achievable sub-goals derived from 012 the instructions and subsequently completes the 013 associated sub-tasks using a pre-trained RL pol-014 icy. We employ the IGLU competition, which 015 is based on the Minecraft-like simulator, as our 016 test environment, and compare our approach 017 to the competition’s top-performing solutions. 018 Our approach outperforms existing solutions in 019 terms of both the quality of the language model 020 and the quality of the structures built within the 021 IGLU environment. 022

Learning Latent Structural Causal Models

Jithendaraa Subramanian

Yashas Annadani

Ivaxi Sheth

Nan Rosemary Ke

Tristan Deleu

Stefan Bauer

Samira Ebrahimi Kahou

Causal learning has long concerned itself with the accurate recovery of underlying causal mechanisms. Such causal modelling enables better e… (see more)xplanations of out-of-distribution data. Prior works on causal learning assume that the high-level causal variables are given. However, in machine learning tasks, one often operates on low-level data like image pixels or high-dimensional vectors. In such settings, the entire Structural Causal Model (SCM) -- structure, parameters, \textit{and} high-level causal variables -- is unobserved and needs to be learnt from low-level data. We treat this problem as Bayesian inference of the latent SCM, given low-level data. For linear Gaussian additive noise SCMs, we present a tractable approximate inference method which performs joint inference over the causal variables, structure and parameters of the latent SCM from random, known interventions. Experiments are performed on synthetic datasets and a causally generated image dataset to demonstrate the efficacy of our approach. We also perform image generation from unseen interventions, thereby verifying out of distribution generalization for the proposed causal model.

2022-10-24

ArXiv (preprint)

Single‐pass stratified importance resampling

Ege Ciklabakkal

Adrien Gruson

Iliyan Georgiev

Toshiya Hachisuka

Resampling is the process of selecting from a set of candidate samples to achieve a distribution (approximately) proportional to a desired t… (see more)arget. Recent work has revisited its application to Monte Carlo integration, yielding powerful and practical importance sampling methods. One drawback of existing resampling methods is that they cannot generate stratified samples. We propose two complementary techniques to achieve efficient stratified resampling. We first introduce bidirectional CDF sampling which yields the same result as conventional inverse CDF sampling but in a single pass over the candidates, without needing to store them, similarly to reservoir sampling. We then order the candidates along a space‐filling curve to ensure that stratified CDF sampling of candidate indices yields stratified samples in the integration domain. We showcase our method on various resampling‐based rendering problems.

2022-07-30

Computer Graphics Forum (published)

Kubric: A scalable dataset generator

Klaus Greff

Francois Belletti

Lucas Beyer

Carl Doersch

Yilun Du

Daniel Duckworth

David J Fleet

Dan Gnanapragasam

Florian Golemo

Charles Herrmann

Thomas Kipf

Abhijit Kundu

Dmitry Lagun

Issam Hadj Laradji

Hsueh-Ti Liu

Henning Meyer

Yishu Miao

Cengiz Oztireli

Etienne Pot … (see 14 more)

Noha Radwan

Daniel Rebain

Sara Sabour

Mehdi S. M. Sajjadi

Matan Sela

Vincent Sitzmann

Austin Stone

Deqing Sun

Suhani Vora

Ziyu Wang

Tianhao Wu

Kwang Moo Yi

Fangcheng Zhong

Andrea Tagliasacchi

Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance o… (see more)f a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises additional privacy, fairness and legal concerns. Synthetic data is a powerful tool with the potential to address these shortcomings: 1) it is cheap 2) supports rich ground-truth annotations 3) offers full control over data and 4) can circumvent or mitigate problems regarding bias, privacy and licensing. Unfortunately, software tools for effective data generation are less mature than those for architecture design and training, which leads to fragmented generation efforts. To address these problems we introduce Kubric, an open-source Python framework that interfaces with PyBullet and Blender to generate photo-realistic scenes, with rich annotations, and seamlessly scales to large jobs distributed over thousands of machines, and generating TBs of data. We demonstrate the effectiveness of Kubric by presenting a series of 13 different generated datasets for tasks ranging from studying 3D NeRF models to optical flow estimation. We release Kubric, the used assets, all of the generation code, as well as the rendered datasets for reuse and modification.

2022-06-18

2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (published)

arxiv.org

Kubric: A scalable dataset generator

Klaus Greff

Francois Belletti

Lucas Beyer

Carl Doersch

Yilun Du

Daniel Duckworth

David J. Fleet

Dan Gnanapragasam

Florian Golemo

Charles Herrmann

Thomas N. Kipf

Abhijit Kundu

Dmitry Lagun

Issam Hadj Laradji

Hsueh-Ti Liu

H. Meyer

Yishu Miao

Cengiz Oztireli

Etienne Pot … (see 14 more)

Noha Radwan

Daniel Rebain

Sara Sabour

Mehdi S. M. Sajjadi

Matan Sela

Vincent Sitzmann

Austin Stone

Deqing Sun

Suhani Vora

Ziyu Wang

Tianhao Wu

Kwang Moo Yi

Fangcheng Zhong

Andrea Tagliasacchi

2022-03-07

ArXiv (preprint)

arxiv.org

Learning to Guide and to Be Guided in the Architect-Builder Problem

Tristan Karch

Clément Moulin-Frier

We are interested in interactive agents that learn to coordinate, namely, a …

2022-01-28

ICLR.cc/2022/Conference (poster)

Attention-based Neural Cellular Automata

Mattie Tesfaldet

Chris Pal

Recent extensions of Cellular Automata (CA) have incorporated key ideas from modern deep learning, dramatically extending their capabilities… (see more) and catalyzing a new family of Neural Cellular Automata (NCA) techniques. Inspired by Transformer-based architectures, our work presents a new class of _attention-based_ NCAs formed using a spatially localized—yet globally organized—self-attention scheme. We introduce an instance of this class named _Vision Transformer Cellular Automata (ViTCA)_. We present quantitative and qualitative results on denoising autoencoding across six benchmark datasets, comparing ViTCA to a U-Net, a U-Net-based CA baseline (UNetCA), and a Vision Transformer (ViT). When comparing across architectures configured to similar parameter complexity, ViTCA architectures yield superior performance across all benchmarks and for nearly every evaluation metric. We present an ablation study on various architectural configurations of ViTCA, an analysis of its effect on cell states, and an investigation on its inductive biases. Finally, we examine its learned representations via linear probes on its converged cell state hidden representations, yielding, on average, superior results when compared to our U-Net, ViT, and UNetCA baselines.

Challenges in leveraging GANs for few-shot data augmentation

Christopher Beckham

Issam Hadj Laradji

Pau Rodriguez

David Vazquez

Chris Pal

2022-01-01

arXiv.org (preprint)