Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Model-completion methods learn a full causal generative model consistent with observational data and a given causal graph, and answer interventional queries via probabilistic inference. We empirically compare two approaches presented recently. One approach learns the model using EM, named EM for Causal Inference (EM4CI). The other approach uses neural networks for completion, yielding two neural causal model approaches, MLE–NCM and GAN–NCM. We evaluate these methods on synthetic discrete benchmarks spanning multiple graph families and scales. Results show that EM4CI seems superior on large graphs in terms of accuracy, while NCM-based methods can be competitive on small models but incur substantially higher computational cost.more » « lessFree, publicly-accessible full text available May 5, 2027
-
Model-completion methods learn a full causal generative model consistent with observational data and a given causal graph, and answer interventional queries via probabilistic inference. We empirically compare two approaches presented recently. One approach learns the model using EM, named EM for Causal Inference (EM4CI). The other approach uses neural networks for completion, yielding two neural causal model approaches, MLE–NCM and GAN–NCM. We evaluate these methods on synthetic discrete benchmarks spanning multiple graph families and scales. Results show that EM4CI seems superior on large graphs in terms of accuracy, while NCM-based methods can be competitive on small models but incur substantially higher computational cost.more » « lessFree, publicly-accessible full text available May 5, 2027
-
Free, publicly-accessible full text available November 1, 2027
-
Current PEFT methods for LLMs can achieve either high quality, efficient training, or scalable serving, but not all three simultaneously. To address this limitation, we investigate sparse fine-tuning and observe a remarkable improvement in generalization ability. Utilizing this key insight, we propose a family of \underline{S}tructured \underline{S}parse \underline{F}ine-\underline{T}uning (\textbf{\model}) methods for LLMs, which \textit{concurrently achieve state-of-the-art fine-tuning performance, training efficiency, and inference scalability}. \model \mbox{accomplishes this by ``selecting sparsely and computing densely". It selects a few} heads and channels in the MHA and FFN modules for each Transformer block, respectively. Next, it co-permutes weight matrices on both sides of the coupled structures in LLMs to connect the selected components in each layer into a dense submatrix. Finally, \model performs in-place gradient updates on all submatrices. Through theoretical analysis and empirical results, our method prevents overfitting and forgetting, delivers SOTA performance on both commonsense and arithmetic reasoning with 4.6$$\%$$ and 1.3$$\%$$ average improvements compared to LoRA, and surpasses full FT by 11.5$$\%$$ when generalizing to various domains after instruction tuning. Using our partial backpropagation algorithm, \model saves training memory up to 3$$\times$$ and improves latency by 1.5-2.7$$\times$$ compared to full FT, while delivering an average 10\% improvement over LoRA on both metrics. We further demonstrate that the weight updates in \model can be decoupled into adapters, enabling effective fusion, fast switch, and efficient parallelism for serving multiple fine-tuned models.more » « less
-
Wang, H; Xiao, X (Ed.)Differential privacy (DP) is applied when fine-tuning pre-trained language models (LMs) to limit leakage of training examples. While most DP research has focused on improving a model’s privacy-utility tradeoff, some find that DP can be unfair to or biased against underrepresented groups. In this work, we extensively analyze the impact of DP on bias in LMs. We find differentially private training can increase the model bias against protected groups w.r.t AUC-based bias metrics. DP makes it more difficult for the model to differentiate between the positive and negative examples from the protected groups and other groups in the rest of the population. Our results also show that the impact of DP on bias is affected by both the privacy protection level and the underlying distribution of the dataset.more » « less
-
Abstract At the end of its mission, the MESSENGER spacecraft's orbit intersected Mercury's nightside magnetic equator at low altitudes below 420 km, enabling the first in situ observations of this region, where the magnetic field strength is typically sub‐dipolar. We present 5 events from these orbits where MESSENGER encountered Earth‐like dipolarization regions characterized by enhanced field strengths up to 20 nT above the intrinsic planetary field, and an average decrease and increase in plasma proton density and temperature, respectively, for 1–2 min periods, comparable to Hermean substorm timescales. The events span local times of 1.5 hr pre‐ and post‐midnight, and are present from the magnetic equator up to magnetic latitudes of at least north. Supported by estimates of decreased flux tube entropy during these events, we suggest these dipolarization regions are formed by the pileup of dipolarization fronts and formation of a substorm current wedge.more » « lessFree, publicly-accessible full text available December 28, 2026
An official website of the United States government

Full Text Available