Loading Events

« All Events

Thesis defense – Yannis Bendi-Ouis

Wednesday 2 December / 13:30

Venue: Centre Broca
Thesis defended in english


IMN
Team: Computational Neurosciences
Thesis directed by Xavier Hinaut

Title

From Hybrid Reservoir–Transformer Models to State Space Models for Efficient Sequence Processing

Abstract

Transformer architectures have revolutionized sequence processing, particularly in the field of language. Their attention mechanism enables direct access to previous elements of a sequence and supports highly parallelizable training. However, this efficiency during training comes with a cost that grows with context length: attention has quadratic complexity with respect to sequence length and, in autoregressive generation, storing past keys and values also requires an amount of memory that increases with the number of processed tokens. In contrast, recurrent architectures maintain a fixed-size state, making them naturally suited to continuous data processing and to inference with constant cost per time step. Their sequential nature, however, limits the parallelization of training. This thesis investigates how to combine the properties of these two types of architectures: the parallel training of Transformers and the constant-cost inference of recurrent models. To this end, it explores the hybridization of Transformers and Reservoir Computing, a particular approach to recurrent neural networks based on dynamical systems. A first architecture, the Echo State Transformer, replaces attention over all past elements with attention applied to a fixed number of dynamic memory units. This work then leads to a reformulation of linear reservoirs. By directly parameterizing their dynamics using complex eigenvalues, it becomes possible to model different temporal scales while replacing recurrent matrix operations with less costly element-wise operations. This progression brings the Dynamical Memory Transformer, an architecture belonging to the class of State-Space Models and based on several selective recurrent memory units. These units are updated as a function of the input through an adaptive leak rate, while attention is used to organize interactions between them. Since this dynamics remains linear, the evolution of the states can be computed using a parallel associative scan. The model therefore retains the advantages of parallelizable training while maintaining a fixed-size state during inference. Finally, the thesis introduces CogScale, a benchmark of synthetic tasks designed to separately evaluate different cognitive aspects such as retention, selection, associative recall, and information manipulation. Overall, this work aims to design architectures capable of preserving precise access to past information while maintaining a constant inference cost per time step.

Key words

Transformers, State Space Models, Reservoir Computing, Dynamical System, Associative Scan, Sequence Processing

Publications

Jury

More details soon

I subscribe to the newsletter:

Details

Date:
Wednesday 2 December
Time:
13:30
Event Categories:
,