Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Search for: All records

Creators/Authors contains: "Zhang, C"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Free, publicly-accessible full text available May 12, 2027
  2. Free, publicly-accessible full text available April 23, 2027
  3. Large language models (LLMs) achieve good performance on challenging reasoning benchmarks, yet could also make basic reasoning mistakes. This contrasting behavior is puzzling when it comes to understanding the mechanisms behind LLMs' reasoning capabilities. One hypothesis is that the increasingly high and nearly saturated performance on common reasoning benchmarks could be due to the memorization of similar problems. In this paper, we systematically investigate this hypothesis with a quantitative measurement of memorization in reasoning tasks, using a dynamically generated logical reasoning benchmark based on Knights and Knaves (K&K) puzzles. We find that LLMs could interpolate and memorize the training puzzles (achieving near-perfect accuracy) after fine-tuning, yet they struggle with slight variations of these puzzles. On the other hand, we show that while fine-tuning leads to heavy memorization, it also consistently improves generalization performance. Through in-depth analyses with perturbation tests, cross difficulty-level transferability, probing model internals, and fine-tuning with wrong answers, we establish that LLMs develop reasoning skills on K&K puzzles alongside memorization. Finally, our analysis based on a per-sample memorization score sheds light on how LLMs switch between reasoning and memorization when solving logical puzzles. 
    more » « less
    Free, publicly-accessible full text available December 20, 2026
  4. Free, publicly-accessible full text available November 1, 2026
  5. Free, publicly-accessible full text available December 2, 2026
  6. ABSTRACT: Ethyl cellulose (EC) is a biocompatible, renewable, and recyclable material with diverse sources, making it an attractive candidate for industrial applications. Electrospinning has gained significant attention for the production of EC fibers. However, conventional electrospinning methods face challenges such as bead formation, low yield, and the absence of porous internal structures, limiting both the functional performance and scalability. This study presents an optimized approach for producing EC fibers by using a gravity-driven ultrahigh-speed electrospinning (GUHS-ES) system. This system leverages gravity to reshape the Taylor cone morphology during electrospinning, enhancing stability and dramatically increasing throughput. As flow rates increase, the Taylor cone contracts inward, while the tip structure expands and stabilizes, reaching maximum size at ultrahigh flow rates (100−150 mL/h). This unique Taylor cone structure enables a fiber production rate of 24.5 g/h, hundreds of times greater than conventional electrospinning techniques. Another advantage of the GUHS-ES system is its ability to achieve both high diameter uniformity and adjustable porosity. At ultrahigh flow rates, the pore sizes of the EC fibers reached 321 nm. The highly porous structure of EC fibers exhibited an absorption capacity of 56.6 to 110.7 times their weight, exceeding most previously reported oil-absorbing materials and demonstrating high efficacy for rapid waste oil absorption. This green, efficient technology represents a promising advancement for the large-scale production and application of natural polymer fibers with broad implications for sustainable industrial processes. 
    more » « less
  7. Phylogenetic estimation is, and has always been, a complex endeavor. Estimating a phylogenetic tree involves evaluating many possible solutions and possible evolutionary histories that could explain a set of observed data, typically by using a model of evolution. Modern statistical methods involve not just the estimation of a tree, but also solutions to more complex models involving fossil record information and other data sources. Markov Chain Monte Carlo (MCMC) is a leading method for approximating the posterior distribution of parameters in a mathematical model. It is deployed in all Bayesian phylogenetic tree estimation software. While many researchers use MCMC in phylogenetic analyses, interpreting results and diagnosing problems with MCMC remain vexing issues to many biologists. In this manuscript, we will offer an overview of how MCMC is used in Bayesian phylogenetic inference, with a particular emphasis on complex hierarchical models, such as the fossilized birth-death (FBD) model. We will discuss strategies to diagnose common MCMC problems and troubleshoot difficult analyses, in particular convergence issues. We will show how the study design, the choice of models and priors, but also technical features of the inference tools themselves can all be adjusted to obtain the best results. Finally, we will also discuss the unique challenges created by the incorporation of fossil information in phylogenetic inference, and present tips to address them. 
    more » « less