scieee AI-readable full text Open interactive document viewer

Learning Light Curve Embeddings with Rotary Masked Autoencoder

Contardo, Gabriella

Abstract

Upcoming data from the Rubin Observatory offer unprecedented opportunities for astronomical data analysis, but also methodological challenges: how can we automatically detect new (sub)classes of events? Can we detect rare or anomalous events? We investigate here the Rotary Masked Auto-Encoder (RoMAE), a recent development of Transformers to process irregularly-sampled multivariate time-series. Self-supervised pretraining has been shown to significantly improve downstream tasks, suggesting that these models can extract and encode high-level information in their embeddings. However, it remains unclear how (and which) information can be retrieved and disentangled in a setup like Rubin’s data. Thus, we explore the properties of RoMAE’s embeddings in different synthetic scenarios using ELASTiCC.v2 and investigate the ability of the embeddings to identify different types of transients and anomalies. We also tentatively explore the application of our approach on Rubin’s Data Preview 1.

Full text

Embeddings (CLS-token) of test examples of the seen-classes + UMAP ( ): Surprisingly clean separation of TDE and SNIa (too good?): this is a bit suspicious... No outliers (expected?) Two clusters of TDEs: very likely due to the observation pattern? (2 gaps vs 1) Learning Light Curve Embeddings with Rotary Masked Autoencoder Gabriella Contardo (University of Nova Gorica, SISSA), Andreja Gomboc (UNG), Uros Zivanovic (University of Trieste, SISSA), Alex Razim (UNG), Eduardo Concepcion Castro (UNG), Saptashwa Bhattacharyya (UNG) arXiv: 2505.20535 A Transformer for Irregularly Sampled Multi-Variate Time-Series Class discovery and anomaly detection in Rubin’s light curves → need to transform the light curves into some feature space. RoMAE: a specific positional encoding (“rotary”) of time (and filters) for irregularly sampled multi-variate time-series. State-ofthe-art results on supervised tasks. Are RoMAE’s embeddings learned in a selfsupervised way (no label) useful for anomaly/class detection in Rubin’s data? Unsupervised training with only two classes of ELAsTiCC2 DP1 light-curves are shorter than ELASTiCC2 + different samplings: could explain why they are well separated (+ training with only 2 classes). Does not explain why one lands in + not in the right region? Concerning behavior in ELAsTiCC: there is such a thing as too good results... Small “clusters” in DP1 hard to interpret; neighbors of SN from Freeburn don’t really look like SN... No clear separation in TDE/SNIa using ZTF. UMAP not ideal: future work train for smaller representations directly; train with more classes/variety. Injecting New Classes + Rubin DP1-ECDFS Real data Light-curves for DP1-ECDFS are constructed from DiaObjects and DiaSources with all flagged observations removed, keeping objects with at least 3 detections in 2 filters Train on DP1 only Train on ZTF TDE & SNIa only Some thoughts According to the models, DP1-ECDFS lightcurves look very different from the ELAsTiCC2 classes seen in training (except this one ) Case 1: TDE+EB Case 2: TDE+SNIa