Publications

Motor cortex latent dynamics encode arm movement direction and urgency independently

Andrea Colins Rodriguez

Matt Perich

Lee Miller

Mark D. Humphries

2023-05-26

bioRxiv (prépublication)

Testing Feedforward Neural Networks Training Programs

Houssem Ben Braiek

Foutse Khomh

2023-05-26

ACM Transactions on Software Engineering and Methodology (publié)

An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics

Saba Ahmadi

Aishwarya Agrawal

2023-05-24

ArXiv (prépublication)

Model evaluation for extreme risks

Toby Shevlane

Sebastian Farquhar

Ben Garfinkel

Mary Phuong

Jess Whittlestone

Jade Leung

Daniel Kokotajlo

Nahema A. Marchal

Markus Anderljung

Noam Kolt

Lewis Ho

Divya Siddarth

Shahar Avin

W. Hawkins

Been Kim

Iason Gabriel

Vijay Bolina

Jack Clark

Paul F. Christiano … (voir 1 de plus)

Allan Dafoe

Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further pro… (voir plus)gress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through"dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through"alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.

2023-05-24

ArXiv (prépublication)

Model evaluation for extreme risks

Toby Shevlane

Sebastian Farquhar

Ben Garfinkel

Mary Phuong

Jess Whittlestone

Jade Leung

Daniel Kokotajlo

Nahema A. Marchal

Markus Anderljung

Noam Kolt

Lewis Ho

Divya Siddarth

Shahar Avin

W. Hawkins

Been Kim

Iason Gabriel

Vijay Bolina

Jack Clark

Paul F. Christiano … (voir 1 de plus)

Allan Dafoe

Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further pro… (voir plus)gress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through"dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through"alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.

2023-05-24

ArXiv (prépublication)

Model evaluation for extreme risks

Toby Shevlane

Sebastian Farquhar

Ben Garfinkel

Mary Phuong

Jess Whittlestone

Jade Leung

Daniel Kokotajlo

Nahema A. Marchal

Markus Anderljung

Noam Kolt

Lewis Ho

Divya Siddarth

Shahar Avin

W. Hawkins

Been Kim

Iason Gabriel

Vijay Bolina

Jack Clark

Paul F. Christiano … (voir 1 de plus)

Allan Dafoe

2023-05-24

ArXiv (prépublication)

Model evaluation for extreme risks

Toby Shevlane

Sebastian Farquhar

Ben Garfinkel

Mary Phuong

Jess Whittlestone

Jade Leung

Daniel Kokotajlo

Nahema A. Marchal

Markus Anderljung

Noam Kolt

Lewis Ho

Divya Siddarth

Shahar Avin

W. Hawkins

Been Kim

Iason Gabriel

Vijay Bolina

Jack Clark

Paul F. Christiano … (voir 1 de plus)

Allan Dafoe

Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further pro… (voir plus)gress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through"dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through"alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.

2023-05-24

ArXiv (prépublication)

Model evaluation for extreme risks

Toby Shevlane

Sebastian Farquhar

Ben Garfinkel

Mary Phuong

Jess Whittlestone

Jade Leung

Daniel Kokotajlo

Nahema A. Marchal

Markus Anderljung

Noam Kolt

Lewis Ho

Divya Siddarth

Shahar Avin

W. Hawkins

Been Kim

Iason Gabriel

Vijay Bolina

Jack Clark

Paul F. Christiano … (voir 1 de plus)

Allan Dafoe

Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further pro… (voir plus)gress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through"dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through"alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.

2023-05-24

ArXiv (prépublication)

De novo motor learning creates structure in neural activity space that shapes adaptation

Joanna C. Chang

Matt Perich

Lee Miller

Juan A. Gallego

Claudia Clopath

2023-05-24

bioRxiv (prépublication)

Realistically distributing object placements in synthetic training data improves the performance of vision-based object detection models

Setareh Dabiri

Vasileios Lioutas

Berend Zwartsenberg

Yunpeng Liu

Matthew Niedoba

Xiaoxuan Liang

Dylan Green

Justice Sefas

Jonathan Wilder Lavington

Frank N. Wood

Adam Ścibior

When training object detection models on synthetic data, it is important to make the distribution of synthetic data as close as possible to … (voir plus)the distribution of real data. We investigate specifically the impact of object placement distribution, keeping all other aspects of synthetic data fixed. Our experiment, training a 3D vehicle detection model in CARLA and testing on KITTI, demonstrates a substantial improvement resulting from improving the object placement distribution.

2023-05-24

ArXiv (prépublication)

Think Before You Act: Decision Transformers with Internal Working Memory

Jikun Kang

Romain Laroche

Xingdi Yuan

Adam Trischler

Xue (Steve) Liu

Jie Fu

Large language model (LLM)-based decision-making agents have shown the ability to generalize across multiple tasks. However, their performan… (voir plus)ce relies on massive data and compute. We argue that this inefficiency stems from the forgetting phenomenon, in which a model memorizes its behaviors in parameters throughout training. As a result, training on a new task may deteriorate the model's performance on previous tasks. In contrast to LLMs' implicit memory mechanism, the human brain utilizes distributed memory storage, which helps manage and organize multiple skills efficiently, mitigating the forgetting phenomenon. Thus inspired, we propose an internal working memory module to store, blend, and retrieve information for different downstream tasks. Evaluation results show that the proposed method improves training efficiency and generalization in both Atari games and meta-world object manipulation tasks. Moreover, we demonstrate that memory fine-tuning further enhances the adaptability of the proposed architecture.

2023-05-24

ArXiv (prépublication)