This program supports AI startups at any time of the year. Benefit from cutting-edge resources and tailored support to accelerate your technology's development.
Offered by Mila and the Public Policy Forum, this program is designed to equip policy and decision makers with the tools to navigate the opportunities and risks of AI. The next cohort will be held in French on September 1-2, 2026, at Mila.
Connect with a Mila academic advisor and current student-researchers to learn more about Mila's community and how to join us on August 19, 31 and September 11, 2026.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Tetradic Dynamics of Dyadic Sensorimotor Coordination: A Multiscale EEG Hyperscanning Study
How do leader and follower roles shape the brain mechanisms that support coordinated action between people? This question has direct therape… (see more)utic relevance for conditions such as schizophrenia and autism spectrum disorder, where the capacity for reciprocal social coordination is a defining vulnerability. Here, we propose a tetradic framework and examine sensorimotor coordination in 16 healthy adult pairs using simultaneous dual-brain EEG hyperscanning, a 2 x 2 within-subject design crossing Role (Leader/Follower) and Condition (Mirroring/Matching). Mirroring required resonance with a partner's movement; Matching required its controlled transformation. This contrast was designed to dissociate automatic from controlled coordination processes across roles. Behaviourally, Mirroring produced a reaction-time advantage that was selective to Followers, a finding replicated in a combined cohort, and consistent with role-dependent attention-inhibition gating. At the neural level, sustained (tonic) activity was dominated by Matching-related frontoparietal engagement regardless of role, while time-resolved (phasic) activity revealed a Mirroring-dominant reorganization that differentiated into role-specific patterns: anticipatory gating in Leaders and post-response inhibitory rebound in Followers. Information-theoretic decomposition of inter-brain coupling identified a Leader-specific predictive signal in medial prefrontal and cingulate cortices, a Follower-specific adaptive signal across sensorimotor and temporal regions, and a shared redundancy scaffold in orbitofrontal and insular cortices. These findings characterize tetradic coordination within dyads as a multiscale, role-asymmetric architecture in which top-down predictive control and bottom-up adaptive regulation are functionally dissociable. The tetradic framework provides an organizing scaffold for this dissociation, and the role-specific signatures it reveals offer candidate biomarkers for clinical populations in whom interpersonal coordination is disrupted.
The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, manual oversight is impr… (see more)actical, resulting in sizeable datasets that can be very noisy. Attempts to mitigate this obstacle to producing performant vision-language models have so far involved heuristics, curated reference datasets, and using pre-trained models. Here we propose a novel, bootstrapped method in which a CLIP model is trained on an evolving, self-selected dataset. This evolving dataset constitutes a balance of filtered, highly probable clean samples as well as diverse samples from the entire distribution. Our proposed Self-Filtering method iterates between training the model and selecting a subsequently improved data mixture. Training on vision-language datasets filtered by the proposed approach improves downstream performance without the need for additional data or pre-trained models.
2026-06-23
Transactions on Machine Learning Research (accepted)
Lacuna is a research map for machine learning that uses LLMs to turn papers and scholarly metadata into markdown summaries, concept elements… (see more), research directions, and research proposals. Each item keeps links to the primary source records and papers that support it. We release the map with web, markdown, and MCP interfaces. Across LitSearch, Multi-XScience-CS/ML, and ScholarQA-CS-ML, Lacuna outperforms OpenScholar with the strongest gains on LitSearch retrieval (Recall@10 0.538 vs. 0.424 for OpenScholar v3). We also evaluate Lacuna Deep Research, a multi-stage report agent over the map, on 25 ReportBench-ML survey tasks: Lacuna Deep Research reaches 0.052 citation F1, 0.339 citation precision, 99 expert-reference hits, and 7.82/10 RACE report quality, while GPT-Researcher reaches 0.039 F1, 0.290 precision, 72 hits, and 5.24/10 RACE.
Con Moto: Embodied Steering of Music Transformers for Live Dance Improvisation
Zhixing Chen
Heidi Lei
Cheng-Zhi Anna Huang
Con Moto is a real-time generative music system for dance improvisation that supports embodied steering of a transformer model with configur… (see more)able levels of agency. While existing frameworks demonstrate the potential of embodied music-making and movement sonification in live performance, achieving both high musical coherence and low-latency responsiveness remains an ongoing challenge. In response, we leverage the musical coherence of real-time MIDI-based transformer models to design an integrated system that translates camera motion data into movement parameters, which in turn control the musical output. Con Moto employs two layers of control strategies: 1) inference-time steering of the transformer model and 2) post-generation rendering using Max/MSP as a control interface and Ableton Live for sound synthesis. We present the system through a duet performance for a live audience, supplemented by qualitative reflections from the dancers and the audience. By reconfiguring how the dancers' movements map to musical functions, we create a system with configurable agency. Fine-grained control over an individual musical voice invited dancers to experience the system as a playable instrument, while abstract musical mappings to a genre's energy opened space for the system to act as an autonomous creative partner. By navigating the aesthetic friction between human intent and AI agency, we explore a dynamic that facilitates a deep, bidirectional feedback loop.
Enhancing Expressive Musical Conversation in the jam_bot
Lancelot Blanchard
Perry Naseck
Katherine Liang
Joel Tan
Heidi Lei
Cheng-Zhi Anna Huang
Joseph Paradiso
Previous work introduced the jam_bot, a real-time system that embeds live music language models capable of generating symbolic music sequenc… (see more)es coherent with a performer's input. The system supports multiple interaction strategies that have been demonstrated in several public performances. However, these strategies limit expressive musical conversation by constraining tempo, form, or musical roles. We extend the jam_bot to support more expressive, open-ended interaction through four key improvements: (1) modeling velocity, a key dimension of expression in symbolic music; (2) increasing model throughput via a ggml implementation–required to accommodate the longer sequences induced by velocity modeling; (3) developing a new training modality that enables free-form call-and-response interaction across varying tempi; and (4) compensating for external MIDI output latency to ensure rhythmic coherence with the performer. We quantitatively evaluate the model throughput improvement and our latency compensation strategy, and offer MIDI samples online. Together, these enhancements enable the jam_bot to engage in natural, expressive musical conversation, eliminating key musical limitations to enable the development of future performances and installations.
Abstract Ceftiofur resistance in canine clinical Escherichia coli is usually associated with extended-spectrum β-lactamases (ESBLs) or AmpC… (see more) β-lactamases. However, some isolates display elevated ceftiofur minimum inhibitory concentrations (MICs) without carrying these known resistance genes. In this study, we identified a phenotype-genotype discordant canine clinical E. coli isolate named 231255, with a ceftiofur MIC of 8 μg/mL. Routine resistance gene screening detected only bla TEM-1 and a chromosomal bla EC variant, bla EC-1149 , which could not adequately explain the elevated ceftiofur MIC. To investigate the underlying mechanism, a genomic library was constructed from genomic DNA of isolate 231255 and screened on ceftiofur-containing plates. Positive clones revealed one candidate determinant: an altered ftsI fragment encoding penicillin-binding protein 3 (PBP3) with a four-amino-acid YRIN insertion downstream of residue P333, previously identified in human E. coli . Previous studies have shown that this type of YRIN/YRIK insertion alone can reduce susceptibility to PBP3-targeting β-lactams, particularly aztreonam and ceftazidime. Functional validation showed that recombinant plasmids carrying the altered ftsI consistently increased the ceftiofur MIC to 4 μg/mL in different recipient backgrounds. These findings provide experimental evidence that ftsI /PBP3 alteration can elevate ceftiofur MIC in canine E. coli . Notably, the isolate belonged to ST410, and phylogenetic analysis indicated that it was not confined to a dog-associated background but instead clustered within a broader lineage shared across multiple sources, highlighting the need for potential dissemination of this mechanism and its associated resistant lineages at the human-companion animal interface.
Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately on underrepresented su… (see more)bgroups. Existing methods address this by adjusting network parameters, guided either by subgroup annotations or inferred pseudo-group labels. Yet at inference, these methods produce only a class prediction, with no insight into a sample's latent subgroup. We propose neural classification trees (NCT), a framework that achieves robustness by encoding subgroup structure in its tree-shaped architecture. By routing each sample to an "easy" or "hard" node of this tree -- based on prediction correctness -- and reusing these routes as pseudo-labels for the next iteration, NCT disentangles conflicting subgroups, without requiring subgroup supervision. We evaluate NCT on five benchmarks spanning binary and multi-class spurious correlations. Our experiments show that the learned tree topology provides strong interpretability by consistently isolating minority subgroups, which provides a transparent mapping between the model architecture and the data's latent group structure, while yielding competitive robustness with state-of-the-art methods.
MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and Ph… (see more)astCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on ∼149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, p 10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI : 10.5281/zenodo.20616818.
Abstract Background Chronic pain is associated with impaired muscle health, but whether these changes reflect site-specific factors, broader… (see more) systemic factors, or both remains unclear. The purpose of this study is to determine whether normative markers of muscle health derived from MRI show site-specific patterns in chronic pain. Methods UK Biobank participants who underwent whole-body MRI from 2006 to 2010 were included in this retrospective cross-sectional study. The MuscleMap Toolbox quantified volume and intramuscular fat (IMF) in 42 muscles of the abdomen, pelvis, and thigh. Normative models trained on a no pain group generated muscle-specific deviations from normal (i.e., Z-scores) for single- and multi-site chronic and acute pain. Results Of 17,843 participants, the primary site-specific analysis included 9,704 no pain, 885 single-site chronic back pain (CBP), 438 single-site chronic hip pain (CHP), and 1,315 single-site chronic knee pain (CKP) participants (n=12,342; mean age 63.7±7.5 years; 52.7% female). Additional analyses included single-site chronic neck/shoulder pain, acute pain, and multi-site chronic pain groups. In CBP, deviations were localized to abdominal muscles, with decreased volume in 6/8 and increased IMF in 6/8. In CHP, deviations were broad, with decreased volume in 3/8 of the abdominal and 14/26 of the thigh muscles, and increased IMF in 6/8 of the abdominal, 5/8 of the pelvic, and 4/26 of the thigh muscles. In CKP, deviations were localized to thigh muscles, with decreased volume in 8/26 and increased IMF in 6/26. Acute pain groups showed no significant differences except for decreased volume in one thigh muscle in acute knee pain. With each additional chronic pain site, volume decreased (β=−.078;IQR:−0.100−0.051), and IMF increased (β=.085;IQR:0.066−0.101). Combined Z-scores classified chronic pain groups better than chance (accuracy: 48.6%;p<.001), but not acute pain groups (accuracy: 39.0%;p=.20). Conclusions Whole-body MRI combined with AI-driven muscle segmentation and normative modeling revealed site-specific patterns of muscle health in single-site chronic pain.
Accurate identification of the epileptogenic zone (EZ) is essential for seizure freedom after resective surgery in drug-resistant epilepsy, … (see more)yet seizure freedom rates remain below 50%. We developed EpiiSLM, a dual foundation model system for EZ identification with stereo-electroencephalography (sEEG), by training a signal foundation model on 104,990 minutes of sEEG recordings from the Montreal Neurological Institute & Hospital, while leveraging all recordings regardless of surgical outcome and anchoring EZ biomarker extraction on non-epileptic signals. A language foundation model then integrates sEEG-derived outputs with multimodal clinical information to produce interpretable predictions. Under leave-one-patient-out evaluation, EpiiSLM achieved 0.978 contact-level positive predictive value (PPV), outperforming the seizure onset zone(SOZ)-as-EZ baseline by 15.1% (p < 0.05), and 100% region-level accuracy; on an external dataset, EpiiSLM achieved 0.857 contact-level PPV. EpiiSLM requires only one night of interictal sleep data, suggesting potential to reduce invasive sEEG monitoring duration and improve surgical outcomes.