Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
This paper investigates the accuracy of generative models and the impact of knowledge transfer on their generation precision. Specifically, we examine a generative model for a target task, fine-tuned using a pre-trained model from a source task. Building on the "Shared Embedding" concept, which bridges the source and target tasks, we introduce a novel framework for transfer learning under distribution metrics such as the Kullback-Leibler divergence. This framework underscores the importance of leveraging inherent similarities between diverse tasks despite their distinct data distributions. Our theory suggests that the shared structures can augment the generation accuracy for a target task, reliant on the capability of a source model to identify shared structures and effective knowledge transfer from source to target learning. To demonstrate the practical utility of this framework, we explore the theoretical implications for two specific generative models: diffusion and normalizing flows. The results show enhanced performance in both models over their non-transfer counterparts, indicating advancements for diffusion models and providing fresh insights into normalizing flows in transfer and non-transfer settings. These results highlight the significant contribution of knowledge transfer in boosting the generation capabilities of these models.more » « lessFree, publicly-accessible full text available December 20, 2027
-
Free, publicly-accessible full text available July 7, 2027
-
Free, publicly-accessible full text available December 1, 2027
-
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal InferenceReliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trustworthy statistical causal inference. Current benchmarks usually involve simplified tasks. For example, these tasks might only ask LLMs to identify semantic causal relationships or draw conclusions directly from raw data. As a result, models may overlook important statistical pitfalls, such as Simpson’s paradox or selection bias. This oversight limits the applicability of LLMs in the real world. To address these limitations, we propose CausalPitfalls, a comprehensive benchmark designed to rigorously evaluate the capability of LLMs in overcoming common causal inference pitfalls. Our benchmark features structured challenges across multiple difficulty levels, each paired with grading rubrics. This approach allows us to quantitatively measure both causal reasoning capabilities and the reliability of LLMs’ responses. We evaluate models using two protocols: (1) direct prompting, which assesses intrinsic causal reasoning, and (2) code-assisted prompting, where models generate executable code for statistical analysis. Additionally, we validate the effectiveness of this judge by comparing its scoring with assessments from human experts. Our results reveal significant limitations in current LLMs when performing statistical causal inference. The CausalPitfalls benchmark provides essential guidance and quantitative metrics to advance the development of trustworthy causal reasoning systems. Our code is publicly available at CausalPitfalls.more » « lessFree, publicly-accessible full text available April 23, 2027
-
Atrial fibrillation (AF), a cardiac arrhythmia characterized by an abnormal and rapid heartbeat, has the potential to develop into stroke, heart failure, and, ultimately, mortality. The electrocardiogram (ECG) is a pivotal tool in the diagnosis of AF, offering a quick, cost‐effective, and non‐invasive mean to record the heart's electrical activity. Recent studies are increasingly engaged in the implementation of deep learning techniques for ECG feature extraction for AF prediction. In addition, the application of Mendelian randomization (MR) methodologies has been investigated to identify causal associations between genetically imputed pre‐defined ECG characteristics and cardiovascular diseases, such as AF. DeepFEIVR, a non‐linear extension of the classical instrumental variable (IV) regression model, was designed with the objective of extracting disease‐associated causal features from high‐dimensional data, such as neuroimaging data. In this article, we applied DeepFEIVR as well as its variant (with residual inclusion), DeepFEIVR‐RI, to the large UK Biobank dataset. The application of DeepFEIVR and DeepFEIVR‐RI showed that the genetic components in ECGs could contribute to the development of AF statistically significantly (pvalues < ). Another contribution of this article is an extension to both DeepFEIVR and DeepFEIVR‐RI to accommodate a large number of IVs. A comparison of results from DeepFEIVR and DeepFEIVR‐RI, based on various choices of IVs, was conducted. Furthermore, we applied a recent algorithm called dnn‐loc, enabling a visual examination on specific ECG components as extracted causal features for AF, thus advancing the understanding of the etiology of AF.more » « less
-
Abstract While whistler‐mode waves are generated by injected anisotropic electrons on the nightside, the observed day‐night asymmetry of wave distributions raises an intriguing question about their generation on the dayside. In this study, we evaluate the distributions of whistler‐mode wave amplitudes and electrons as a function of distance from the magnetopause (MP) on the dayside from 6 to 18 hr in magnetic local time (MLT) within ±18° of magnetic latitude using the Time History of Events and Macroscale Interaction During Substorms measurements from June 2010 to August 2018. Specifically, under different levels of solar wind dynamic pressure and geomagnetic index, we conduct a statistical analysis to examine whistler‐mode wave amplitude, as well as anisotropy and phase space density (PSD) of source electrons across 1–20 keV energies, which potentially provide a source of free energy for wave generation. In coordinates relative to the MP, we find that lower‐band (0.05–0.5fce) waves occur much closer to the MP than upper‐band (0.5–0.8fce) waves, wherefceis electron cyclotron frequency. Our statistical results reveal that strong waves are associated with high anisotropy and high PSD of source electrons near the equator, indicating a preferred region for local wave generation on the dayside. Over 10–14 hr in MLT, as latitude increases, electron anisotropy decreases, while whistler‐mode wave amplitudes increase, suggesting that wave propagation from the equator to higher latitudes, along with amplification along the propagation path, is necessary to explain the observed waves on the dayside.more » « less
-
Transformer models have been widely investigated in different domains by providing long-range dependency handling and global contextual awareness, driving the development of popular AI applications such as ChatGPT, Gemini, and Alexa. State Space Models (SSMs) have emerged as strong contenders in the field of sequential modeling, challenging the dominance of Transformers. SSMs incorporate a selective mechanism that allows for dynamic parameter adjustment based on input data, enhancing their performance. However, this mechanism also comes with increasing computational complexity and bandwidth demands, posing challenges for deployment on resource-constraint mobile devices. To address these challenges without sacrificing the accuracy of the selective mechanism, we propose a sparse learning framework that integrates architecture-aware compiler optimizations. We introduce an end-to-end solution–C 4 n kernel sparsity, which prunes n elements from every four contiguous weights, and develop a compiler-based acceleration solution to ensure execution efficiency for this sparsity on mobile devices. Based on the kernel sparsity, our framework generates optimized sparse models targeting specific sparsity or latency requirements for various model sizes. We further leverage pruned weights to compensate for the remaining weights, enhancing downstream task performance. For practical hardware acceleration, we propose C 4 n -specific optimizations combined with a layout transformation elimination strategy. This approach mitigates inefficiencies arising from fine-grained pruning in linear layers and improves performance across other operations. Experimental results demonstrate that our method achieves superior task performance compared to other semi-structured pruning methods and achieves up-to 7→ speedup compared to llama.cpp framework on mobile devices.more » « less
-
Abstract Electromagnetic ion cyclotron waves in the Earth's outer radiation belt drive rapid electron losses through wave‐particle interactions. The precipitating electron flux can be high in the hundreds of keV energy range, well below the typical minimum resonance energy. One of the proposed explanations relies on nonresonant scattering, which causes pitch‐angle diffusion away from the fundamental cyclotron resonance. Here we propose the fractional sub‐cyclotron resonance, a second‐order nonlinear effect that scatters particles at resonance ordern = 1/2, as an alternate explanation. Using test‐particle simulations, we evaluate the precipitation ratios of sub‐MeV electrons for wave packets with various shapes, amplitudes, and wave normal angles. We show that the nonlinear sub‐cyclotron scattering produces larger ratios than the nonresonant scattering when the wave amplitude reaches sufficiently large values. The ELFIN CubeSats detected several events with precipitation ratio patterns matching our simulation, demonstrating the importance of sub‐cyclotron resonances during intense precipitation events.more » « less
-
Abstract Interchange instability is known to drive fast radial transport of electrons and ions in Jupiter's inner and middle magnetosphere. In this study, we conduct a statistical survey to evaluate the properties of energetic particles and plasma waves during interchange events using Juno data from 2016 to 2023. We present representative examples of interchange events followed by a statistical analysis of the spatial distribution, duration and spatial extent. Our survey indicates that interchange instability is predominant atM‐shells from 6 to 26, peaking near 17 with an average duration of minutes and a correspondingM‐shell width of <∼0.05. During interchange events, the associated plasma waves, such as whistler‐mode, Z‐mode, and electron cyclotron harmonic waves exhibit a distinct preferential location. These findings provide valuable insights into particle transport and the source region of plasma waves in the Jovian magnetosphere, as well as in other magnetized planets within and beyond our solar system.more » « less
An official website of the United States government

Full Text Available