<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>The Ruminococcus bromii amylosome protein Sas6 binds single and double helical α-glucan structures in starch</title></titleStmt>
			<publicationStmt>
				<publisher>Nature Structure and Molecular Biology</publisher>
				<date>01/04/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10489444</idno>
					<idno type="doi">10.1038/s41594-023-01166-6</idno>
					<title level='j'>Nature Structural &amp; Molecular Biology</title>
<idno>1545-9993</idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Amanda L. Photenhauer</author><author>Rosendo C. Villafuerte-Vega</author><author>Filipe M. Cerqueira</author><author>Krista M. Armbruster</author><author>Filip Mareček</author><author>Tiantian Chen</author><author>Zdzislaw Wawrzak</author><author>Jesse B. Hopkins</author><author>Craig W. Vander Kooi</author><author>Štefan Janeček</author><author>Brandon T. Ruotolo</author><author>Nicole M. Koropatkin</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Resistant starch is a prebiotic accessed by gut bacteria with specialized amylases and starch-binding proteins. The human gut symbiont Ruminococcus bromii expresses Sas6 (Starch Adherence System member 6), which consists of two starch-specific carbohydrate-binding modules from family 26 (RbCBM26) and family 74 (RbCBM74). Here, we present the crystal structures of Sas6 and of RbCBM74 bound with a double helical dimer of maltodecaose. The RbCBM74 starch-binding groove complements the double helical α-glucan geometry of amylopectin, suggesting that this module selects this feature in starch granules. Isothermal titration calorimetry and native mass spectrometry demonstrate that RbCBM74 recognizes longer single and double helical α-glucans, while RbCBM26 binds short maltooligosaccharides. Bioinformatic analysis supports the conservation of the amylopectin-targeting platform in CBM74s from resistant-starch degrading bacteria. Our results suggest that RbCBM74 and RbCBM26 within Sas6 recognize discrete aspects of the starch granule, providing molecular insight into how this structure is accommodated by gut bacteria.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>The gut microbiota, the consortium of microbes that resides in the human gastrointestinal tract, influences many aspects of host physiology including digestive health <ref type="bibr">[1]</ref>. The composition of the gut microbiota is modulated by the human diet <ref type="bibr">[2]</ref><ref type="bibr">[3]</ref><ref type="bibr">[4]</ref>. After host nutrient absorption in the small intestine, indigestible dietary fiber transits the large intestine and becomes food for gut microbes <ref type="bibr">[3]</ref>. Bacterial fermentation of dietary carbohydrates produces beneficial short chain fatty acids including butyrate, a primary carbon source for colonocytes that also has systemic antiinflammatory and anti-tumorigenic properties <ref type="bibr">[3,</ref><ref type="bibr">5]</ref>.</p><p>Resistant starch is a prebiotic fiber that tends to increase butyrate in the large intestine <ref type="bibr">[6]</ref>. Starch is a glucose polymer composed of branched, soluble amylopectin and coiled insoluble amylose <ref type="bibr">[7,</ref><ref type="bibr">8]</ref>. Breakdown of starch starts with human salivary and pancreatic amylases which release maltooligosaccharides for absorption in the small intestine <ref type="bibr">[9]</ref>. However, a portion of starch is indigestible by human amylases and is termed resistant starch (RS) <ref type="bibr">[9]</ref>. Raw, uncooked starch granules are resistant to digestion in the upper gastrointestinal tract due to the tight packing of constituent amylose and amylopectin into semi-crystalline, insoluble granules <ref type="bibr">[7]</ref>. This type of resistant starch, called RS2, becomes food for gut bacteria that can adhere to and deconstruct granules, releasing glucose and maltooligosaccharides that cross-feed other organisms <ref type="bibr">[9]</ref>.</p><p>Human gut bacteria that degrade RS2 in vitro include Bifidobacterium adolescentis and Ruminococcus bromii <ref type="bibr">[10]</ref><ref type="bibr">[11]</ref><ref type="bibr">[12]</ref><ref type="bibr">[13]</ref><ref type="bibr">[14]</ref>. R. bromii is a Gram-positive anaerobe that increases in relative abundance in the gut upon host consumption of resistant potato or corn starch <ref type="bibr">[10,</ref><ref type="bibr">15,</ref><ref type="bibr">16]</ref>. R. bromii is a keystone species for RS2 degradation because it cross-feeds butyrate-producing bacteria <ref type="bibr">[10]</ref>. R. bromii synthesizes multi-protein starch-degrading complexes called amylosomes via protein-protein interactions between dockerin and complementary cohesin domains <ref type="bibr">[17]</ref><ref type="bibr">[18]</ref><ref type="bibr">[19]</ref>.</p><p>As many as 32 R. bromii proteins have predicted cohesin or dockerin domains including amylases, pullulanases, starch-binding proteins, and proteins of unknown function <ref type="bibr">[17,</ref><ref type="bibr">20]</ref>. Many have carbohydrate-binding modules (CBMs) that presumably aid in binding starch and tether the bacteria to its food source <ref type="bibr">[21]</ref>.</p><p>CBMs are classified by amino acid sequence into numbered families and include members that bind only soluble starch and some that also bind granular starch <ref type="bibr">[21,</ref><ref type="bibr">22]</ref>. One such family is CBM74 which was discovered as a discrete domain (MaCBM74) of a multimodular amylase from the potato starch-degrading bacterium, Microbacterium aurum <ref type="bibr">[22]</ref>. MaCBM74 binds amylose and amylopectin as well as raw wheat, corn, and potato starch granules <ref type="bibr">[22]</ref>. The CBM74 family is unique as it is ~300 amino acids, two to three times larger than most starch-binding CBMs <ref type="bibr">[21]</ref>.</p><p>CBM74 domains are typically found in multimodular enzymes that include a glycoside hydrolase family 13 (GH13) domain for hydrolyzing starch and are flanked by a starch-binding CBM from family 25 or 26 (CBM25 or CBM26) <ref type="bibr">[21,</ref><ref type="bibr">22]</ref>. Most CBM74 family members are encoded by gut microbes and 70% are found in Bifidobacteria <ref type="bibr">[22]</ref>. The genomes of R. bromii and B. adolescentis each encode one putative CBM74-containing protein. The prevalence of CBM74 domains encoded within the genomes of RS2-degrading bacteria, and its increased representation in metagenomic and metatranscriptomic analyses from host diet studies, suggest a role for this module in RS2 recognition in the distal gut <ref type="bibr">[23]</ref><ref type="bibr">[24]</ref><ref type="bibr">[25]</ref>.</p><p>The R. bromii starch adherence system protein 6 (Sas6) is a secreted protein of 734 amino acids that contains both a CBM26 and CBM74 followed by a C-terminal dockerin type 1 domain <ref type="bibr">[26,</ref><ref type="bibr">27]</ref>. Here we present the biochemical characterization and crystal structure of Sas6, providing the first view of the CBM74 domain and its juxtaposition with the CBM26 domain. The co-crystal structure of RbCBM74 with a double helical dimer of maltodecaose, which mimics the architecture of double helical amylopectin in starch granules, revealed recognition via an elongated groove spanning the domain. RbCBM74 exclusively binds longer maltooligosaccharides (&#8805; 8 glucose units), and native mass spectrometry suggests that both single and double helical &#945;-glucans are recognized, providing flexible recognition of amylose and amylopectin. Our biochemical data demonstrate that CBM26 and CBM74 recognize different &#945;-glucan moieties within starch granules leading to overall enhanced granule binding.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head><p>Modular Architecture of Sas6 -Sas6 consists of five discrete domains: an N-terminal CBM26 (RbCBM26), a CBM74 domain (RbCBM74) flanked by Bacterial Immunoglobulin-like (BIg) domains, and a C-terminal dockerin type I (Fig. <ref type="figure">1A</ref>) <ref type="bibr">[27]</ref>. Sas6 is encoded at the WP_015523730 locus (formerly RBR_14490 or Doc6, UniProt: A0A2N0UYM2) and includes a Gram-positive signal peptide (residues 1-30) that presumably targets the protein for secretion.</p><p>RbCBM74 spans residues 242-572 based on an alignment with annotated CBM74 domains <ref type="bibr">[22]</ref>.</p><p>We used InterProScan to annotate the remaining sequence which added the Bacterial Immunoglobulin-like (BIg, Pfam 02368) domain A (BIgA), but did not predict BIgB, which we identified via structure determination <ref type="bibr">[28]</ref>.</p><p>Sas6 Cell Localization -Though Sas6 has a signal peptide it is unknown whether it is a constituent of a cell-bound amylosome, or part of a freely secreted complex <ref type="bibr">[20]</ref>. R. bromii synthesizes five scaffoldin (Sca) proteins that have cohesins for amylosome assembly; Sca2 and Sca5 are cell-bound and Sca1, Sca3, and Sca4 are freely secreted <ref type="bibr">[20]</ref>. The cognate cohesin for the Sas6 dockerin is unknown. Sas6 is detected in the cell-free supernatant of R. bromii cultures in stationary phase but also elutes from the surface of exponentially growing cells with EDTA which disrupts the calcium-dependent cohesin-dockerin interaction <ref type="bibr">[17,</ref><ref type="bibr">29]</ref>. To determine the localization of Sas6, we grew cells to mid-log phase on potato amylopectin and performed a Western Blot with custom antibodies against recombinant Sas6 (Fig. <ref type="figure">1B</ref>). Sas6 was detected in the cell fraction and not the cell-free culture supernatant (Fig. <ref type="figure">1B</ref>), and was visualized on the cell surface via immunofluorescence (Fig <ref type="figure">1C</ref>). Therefore, we conclude that Sas6 is a component of a cell-surface amylosome in actively growing cells. It is possible that Sas6 localization is dependent upon growth phase, as are cellulosome components in some organisms, explaining its previous detection in culture supernatant <ref type="bibr">[17]</ref>. Alternatively, R. bromii, like some cellulosomeproducing bacteria, may release cell-surface amylosomes in stationary phase <ref type="bibr">[30]</ref>.</p><p>Sas6 Starch Binding -CBM26 and CBM74 are putative raw starch-binding families <ref type="bibr">[22,</ref><ref type="bibr">31]</ref>. Plant sources of granular starch differ greatly in granule organization, including crystallinity (e.g., packing of the long helical chains), length of &#945;1,4-linked chains, amylose location and organization, water content, and trace elements <ref type="bibr">[7]</ref>. We used a truncated construct of Sas6 (residues 31-665) lacking the C-terminal dockerin domain, herein called Sas6T, to test Sas6 binding to starch polysaccharides. Sas6T binds potato, corn, and wheat starch granules, with the highest fraction of protein bound to corn starch, and no non-specific binding to Avicel (crystalline cellulose) (Fig. <ref type="figure">1D</ref>). Of note, corn starch has a smaller granule size and therefore a larger surface area to mass ratio <ref type="bibr">[8]</ref>. We tested Sas6T binding to amylopectin and amylose, as well as glycogen and pullulan via affinity PAGE. Glycogen is similar to amylopectin with more frequent &#945;1,6 branching (every 6-15 residues for liver glycogen compared to 15-25 residues for amylopectin) <ref type="bibr">[32,</ref><ref type="bibr">33]</ref>. Pullulan is a fungal &#945;-glucan composed of repeating &#945;1,6-linked maltotriose units <ref type="bibr">[34]</ref>.</p><p>Sas6T binds amylose, amylopectin (potato and corn), and glycogen but has less affinity for pullulan suggesting a preference for longer &#945;1,4-linked regions within the polysaccharide (Fig. <ref type="figure">1E</ref>). Sas6T does not bind dextran, a bacterially derived exopolysaccharide of &#945;1,6-linked glucose <ref type="bibr">[35]</ref>, demonstrating its specificity for starch.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structure of Sas6 -</head><p>The structure of Sas6T with &#945;-cyclodextrin (ACX), was determined via single-wavelength anomalous dispersion of intrinsic sulfur-containing residues to a resolution of 1.6&#197; (Rwork=16.8%, Rfree=21.2%) (Table <ref type="table">1</ref>). The final model contained two molecules of Sas6T in the asymmetric unit, with four Ca 2+ per chain and one molecule of ACX bound at the RbCBM26 domain. The Sas6T structure determined with ACX was used to phase a dataset from unliganded crystals (2.2&#197;, Rwork=19.7%, Rfree=25.5%) (Table <ref type="table">1</ref>). The overall crystal structure of Sas6T is compact, with RbCBM26, BIgA and BIgB forming an arc over RbCBM74 (Fig. <ref type="figure">2A</ref>).</p><p>RbCBM26, RbCBM74, and the dockerin domain are separated by BIgA (light grey) and BIgB (dark grey), respectively (Fig. <ref type="figure">2A</ref>, Extended Data Fig <ref type="figure">1A</ref>). Ig-like or fibronectin-III domains act as spacers in multi-modular glycoside hydrolases including GH13s that target starch <ref type="bibr">[36]</ref>.</p><p>BIgA and BIgB interact via hydrogen bonding with 354&#197; of buried surface area <ref type="bibr">[37]</ref> (Extended Data Fig <ref type="figure">1B</ref>). This interaction may help stabilize or orient the CBM74 domain or the BIgs may act as a hinge between the CBMs. The two chains in the asymmetric unit exhibit some flexibility resulting in different positioning between the RbCBM26 binding site and the RbCBM74 domain (Fig. <ref type="figure">2B</ref>).</p><p>Small Angle X-Ray Scattering -To better connect how our crystal structures correlate with conformational flexibility in solution, we used size-exclusion chromatography coupled with small angle x-ray scattering (SEC-SAXS) on Sas6T (Table <ref type="table">1</ref>). The elution separated out several peaks, including a single strong peak for that was well separated and monodisperse as indicated by the constant radius of gyration (Rg) across the eluted peak (Extended Data Fig <ref type="figure">2A</ref>). The Guinier fit of a subtracted scattering profile created from that peak gave Rg and I(0) values of 29.44 &#177; 0.04&#197; and 0.04 &#177; 3.65 x 10 -5 and the fit and normalized fit residuals confirmed this peak was monodisperse (Extended Data Fig <ref type="figure">2B</ref>). The molecular weight of Sas6T from the SAXS data was calculated to be 61.0 kDa (theoretical 68.9 kDa) indicating it is primarily monomeric in solution <ref type="bibr">[38]</ref>. The Dmax from the P(r) function for Sas6T is 90&#197;. The overall shape of the P(r) function for Sas6T, calculated by indirect Fourier transform (IFT) using GNOM, has a relatively Gaussian shape that is characteristic of a globular compact particle with the main peak at r = ~30 &#197; (Extended Data Fig <ref type="figure">2C</ref>) <ref type="bibr">[39]</ref>. There is a small peak at r = 55&#197; which suggests there are two structurally separate motifs, possibly RbCBM26 and RbCBM74. The dimensionless Kratky plot maxima for Sas6T are typical for a rigid globular protein (Extended Data Fig <ref type="figure">2D</ref>). The small plateau in the mid to high q region, around qRg = 5 in the dimensionless Kratky plot indicates some extension or disorder in the system. These results suggest the presence of two separate modules with flexibility between them, likely corresponding to the two CBMs.</p><p>We tested whether the crystal structure matched the solution data by fitting the crystal structure to the SAXS data using FoXS <ref type="bibr">[40]</ref>. The fit had a &#967; 2 = 2.46 and showed systematic deviations in the normalized fit residual (Extended Data Fig <ref type="figure">2E</ref>). This highlights that there are significant differences between the lowest energy conformation of Sas6T in the crystal structure and the structure of Sas6T in solution. We then used MultiFoXS with our high-resolution structure of Sas6T to account for the flexibility, assigning the linkers between the domains (residues 130-137 and 572-583) as flexible <ref type="bibr">[40]</ref>. MultiFoXS gave a best fit with a 1-state solution with a &#967; 2 = 0.96 and calculated Rg of 29.2&#197; which corroborates the Guinier Rg calculation (Extended Data  <ref type="bibr">[43,</ref><ref type="bibr">44]</ref>. Ca 2+ -1 and Ca 2+ -2 are separated by 3.8&#197; and share three coordinating residues but only Ca 2+ -2 is surface exposed. Ca 2+ -3 is abutted by the loop connecting &#946;-strands 2 and 3 and Ca 2+ -4 is at the center of a loop formed by residues 256-264 and conserved with TmCBM9.</p><p>Like TmCBM9, the Ca 2+ ions in the RbCBM74 structure may be important for structural stability <ref type="bibr">[45]</ref>.</p><p>Molecular Basis of RbCBM26 Binding -The N-terminal RbCBM26 displays a &#946;-sandwich consistent with other members of the CBM26 family <ref type="bibr">[21]</ref>. In both chains of the asymmetric unit, CH/&#960; stacking with ACX is provided by W63 and Y55 with hydrogen bonding mediated by Y53, K101, Q103, and the peptidic oxygen of A107 (Fig. <ref type="figure">2E</ref>). In chain A only, K97 provides hydrogen bonding with O3 of Glc6. In chain B, ACX lies 3.2&#197; from S286 of the CBM74 domain and hydrogen bonds with O2 and O3 of Glc3. In contrast, S286 is 9.5&#197; from ACX in chain A. The top structural homologs of RbCBM26 from DALI are the CBM25 from Bacillus halodurans C-125 (BhCBM26) from &#945;-amylase G-6 (PDB ID: 2C3V-A, Z-score: 12.4, RMSD 1.9&#197;, identity: 16%) and CBM26 (BhCBM26) from the same enzyme (PDB ID: 6B3P-B, Z-score: 12.1, RMSD 1.9&#197;, identity: 20%) <ref type="bibr">[41,</ref><ref type="bibr">46]</ref>. Another top DALI result is ErCBM26b of Amy13K from Eubacterium rectale (PDB ID 2C3H-B, Z-score: 10.8, RMSD 1.7&#197;, identity: 19%). In all three CBM26 structures, the structure and aromatic platforms for ligand recognition are conserved (Extended Data Fig <ref type="figure">4AB</ref>).</p><p>RbCBM26, in contrast to ErCBM26 and BhCBM26, has a longer loop containing K97 and K101 that provide additional hydrogen bonding with ACX. Unlike BhCBM26, RbCBM26 does not undergo a conformational change upon ligand binding (Extended Data Fig 4C) <ref type="bibr">[31]</ref>. A sequence alignment with CBM26 members BhCBM26, ErCBM26 and the Lactobacillus amylovorus &#945;amylase CBM26 (LaCBM26), demonstrates conservation of the aromatic platform but more variation in the hydrogen-bonding network (Extended Data Fig <ref type="figure">4A</ref>). Sas6 W63 corresponds to LaCBM26 W32 that, when mutated, results in complete loss of binding <ref type="bibr">[47]</ref>. The R. bromii protein Sas20 has a CBM26-like domain that shares 26% sequence identity with RbCBM26, yet RbCBM26 shares more structural similarity with BhCBM26 and ErCBM26 <ref type="bibr">[29]</ref>.</p><p>Binding Mechanism of Sas6 -We expressed the individual Sas6 CBMs and included the BIgA/B domains with the CBM74 (BIg-RbCBM74-BIg, residues 134-665) to enhance solubility.</p><p>Sas6T and BIg-RbCBM74-BIg bound granular corn and potato starch, but RbCBM26 did not bind either insoluble starch at detectable levels (Fig. <ref type="figure">2F</ref>). Sas6T binds to more of the corn starch granule, (Kd = 2.8&#181;M &#177; 0.4, Bmax = 0.21&#181;mol/g &#177; 0.01) but has a modestly higher affinity for potato starch (Kd = 1.9&#181;M &#177; 0.3, Bmax = 0.030&#181;mol/g &#177; 0.001), which might be a function of the smaller granule size and larger surface to mass ratio for corn starch. Exclusion of the RbCBM26 in the BIg-RbCBM74-BIg construct led to slightly better binding to corn starch (Kd = 1.5&#181;M &#177; 0.3, Bmax = 0.18&#181;mol/g &#177; 0.008) and modestly higher affinity but less overall binding to potato starch (Kd = 0.51&#181;M &#177; 0.13, Bmax = 0.015&#181;mol/g &#177; 0.001). The saturation curve for BIg-RbCBM74-BIg closely resembles that of Sas6T and there is minimal binding by RbCBM26, suggesting that RbCBM74 drives insoluble starch binding.</p><p>The molecular patterns on the surface of starch granules differs between plant sources and remains an active area of research <ref type="bibr">[48]</ref><ref type="bibr">[49]</ref><ref type="bibr">[50]</ref><ref type="bibr">[51]</ref>. The "hairy billiard ball model" to describe starch granules postulates that the granule surface has block-like clusters of amylopectin chains with hair-like extensions of amylose penetrating through the amylopectin <ref type="bibr">[50]</ref>. Sas6T and BIg-RbCBM74-BIg bind amylose and amylopectin whereas RbCBM26 only binds to amylopectin with apparently low affinity based upon the relatively small change in migration (Fig. <ref type="figure">2G</ref>). This suggests that RbCBM74 drives binding of Sas6 to the long, tightly packed helices of amylose at the surface of the starch granule.</p><p>Using isothermal titration calorimetry (ITC), we found that Sas6T and BIg-RbCBM74-BIg bound amylopectin with sub-micromolar affinity whereas binding was not detectable for RbCBM26 (Table <ref type="table">2</ref>; Extended Data Fig 5A) <ref type="bibr">[52]</ref>. Sas6T binds maltotriose (G3), maltoheptaose (G7), maltooctaose (G8) with a Kd in the hundreds of &#181;M but exhibits a Kd of ~5&#181;M for maltodecaose (G10) (Table <ref type="table">2</ref>; Extended Data Fig <ref type="figure">5B</ref>). Interestingly, RbCBM26 binds shorter linear oligosaccharides (G3, G7) and cyclodextrins, while BIg-RbCBM74-BIg had no detectable affinity for these sugars (Table <ref type="table">2;</ref><ref type="table"/>  linkages are not specifically recognized by either domain. We determined that BIg-RbCBM74-BIg binds exclusively longer &#945;-glucans of at least 8 residues. Notably, &#945;1,4-linked glucose polymers form double helices at 10 glucose units due to internal hydrogen bonding so we hypothesized that RbCBM74 might accommodate starch helices <ref type="bibr">[8]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Molecular Basis of RbCBM74</head><p>Binding -We co-crystallized BIg-RbCBM74-BIg with maltodecaose (G10) to 1.70&#197; resolution (Rwork=17.9%, Rfree=19.9%) (Fig. <ref type="figure">3A</ref>). Remarkably, we observed two molecules of G10 as an extended double helix of ~42&#197; along the face of RbCBM74 extending from S286 (reducing ends) to W373 (non-reducing ends). There was strong electron density for 12 glucoses in one molecule, and nine glucoses in the other chain, likely reflecting varied occupancy of the helix along the binding cleft (Fig. <ref type="figure">3B</ref>). H289, F326, and W373 stood out as surface exposed aromatic residues that might be providing CH-&#960; mediated stacking (Fig. <ref type="figure">3C</ref>). Canonical starch-binding domains feature two or three aromatic residues for pi-stacking interactions with the aglycone face of maltooligosaccharides, but RbCBM74 is designed for extensive hydrogen-bonding interactions with longer oligosaccharides and starch <ref type="bibr">[21]</ref>. The binding site is continuous and each G10 molecule interacts with protein as a stretch of three Glcs at a time, before the natural helical curvature brings the chain out of the contact with the protein (Fig. <ref type="figure">3D</ref>). For example, at the non-reducing end, Glc 1-3 of G10 chain A (G10A) fit into the ligandbinding groove, while Glcs 4-6 of G10A are solvent exposed and Glc 1-3 of G10 chain B (G10B) then fill the cavity. Along the length of the cavity, from the non-reducing end to the reducing end, Glcs 1-3 and 7-9 of both G10A and G10B alternate to fill this binding site.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>An overlay of the</head><p>The binding cleft features a network of residues that hydrogen bond to the hydroxyl groups of glucose (Fig. <ref type="figure">3E</ref>). At the non-reducing end, Glc A1 hydrogen bonds with the indole nitrogen of W373. Glc A2 stacks with W373 with hydrogen bonding provided by G374 and N403. Glc A3 hydrogen bonds with S338. The other molecule of G10 (B) contacts the next part of the binding groove and is anchored by hydrogen bonding of Glc B3 by R336 and Y524. Where the first molecule turns back into the binding groove, Glc A8 hydrogen bonds with E290, D549, and K556. Glc A9 hydrogen bonds with the backbone of H289 and pi stacks with F326. The H289 side chain hydrogen bonds with Glc B7 and provides aromatic character for pi stacking with Glc B8. Near the region of RbCBM74 that lies adjacent to RbCBM26, K464 and S286 hydrogen bond with Glc B9.</p><p>To define the starch-binding properties of RbCBM74 in solution, we employed Hydrogen-Deuterium eXchange Mass Spectrometry (HDX-MS). The conformational dynamics of BIg-RbCBM74-BIg alone and in the presence of G10 were measured over a 4-log timescale (Extended Data Fig <ref type="figure">7AB</ref>). The overall conformational dynamics of the apo protein were consistent with the determined crystal structure, in terms of well-ordered domains and associated loops or flexible regions. The flanking BIg domains showed higher exchange rates than the core CBM74 domain. Intriguingly, the linker regions between domains do not show differentially high dynamic exchange, as would be expected for flexibly tethered independent domains, further supporting the integral nature of BIg-RbCBM74-BIg motif.</p><p>The binding of G10 to RbCBM74 was explored by differential protection from exchange in the absence and presence of G10. Significant protection was observed in the presence of G10, while no significant increases in exchange were observed (Extended Data Fig <ref type="figure">7C</ref>). This is consistent with the minimal global conformation changes between the two states of the protein. The protected regions upon G10 binding were highly localized to a single surface binding region (Fig. <ref type="figure">3F</ref>). This protected region constitutes a single extended surface, which directly overlaps with the G10 binding site observed in the co-crystal structure (Fig. <ref type="figure">3EF</ref>). With the exception of the peptide from A314-Y318 (ANTTY), each of the protected peptides identified by HDX-MS contains at least one key binding residue identified from the co-crystal structure (Fig. <ref type="figure">3E</ref>). These data provide a comprehensive picture of the structural dynamics of RbCBM74 binding to long maltooligosaccharides via an extended starch binding cleft.</p><p>RbCBM74 Mutational Studies -Because most CBM binding is mediated by aromatics, we hypothesized that mutation of W373, F326, or H289 to Ala would dramatically decrease or eliminate binding. We tested maximum binding of each of the aromatic mutants to insoluble corn (1%) and potato starch (5%). The W373A and H289A constructs lost the ability to bind to insoluble corn starch while binding of the F326A construct was greatly reduced (Fig. <ref type="figure">4A</ref>). This trend was somewhat different for potato starch, in which a lower percentage of H289A bound compared to the F326A and W373A mutants. By affinity PAGE, neither the W373A nor the F326A mutant lost appreciable binding to amylopectin while the H289A mutant had a modest decrease in binding to potato amylopectin (Fig. <ref type="figure">4B</ref>). When we quantified binding via ITC, W373A lost all binding for G10 while H289A and F326A had a ~10-20-fold decrease in affinity (Table <ref type="table">2</ref>, Extended Data Fig <ref type="figure">8A</ref>). On potato amylopectin, F326A had a 10-fold reduction in affinity while H289A and W373A exhibited a ~20-fold reduction (Extended Data Fig <ref type="figure">8B</ref>). That single mutations do not eliminate binding is perhaps not surprising given the extensive binding platform. Moreover, the enhanced affinity of these mutants to amylopectin over G10 further suggests that productive interactions with the protein extend beyond a 10-glucose unit footprint. Indeed, the somewhat staggered double helical G10 bound in our crystal structure suggests that at least 12 glucose units contribute to binding (Fig. <ref type="figure">3D</ref>).</p><p>Native mass spectrometry -ITC revealed a binding stoichiometry of 1:1 between BIg-RbCBM74-BIg and G10, while the co-crystal structure demonstrates that two molecules of G10 are accommodated. To better determine the stoichiometry of this binding event, we employed native mass spectrometry in the presence of varying concentrations of G10 (Fig. <ref type="figure">5A</ref>). Each observed state differed by ~1639 Da, which agrees with the theoretical mass of G10 (Extended Data Table <ref type="table">1A</ref>). To obtain binding affinities, we summed the peak intensities of all abundant charge states in our mass spectra and analyzed these intensity values as described previously <ref type="bibr">[53]</ref> (Extended Data Table <ref type="table">1B</ref>). The Kd for BIg-RbCBM74-BIg was determined to be 3.8 &#177; 0.5&#181;M, which agrees with our ITC data. As the concentration of ligand is increased, ligand molecules can bind nonspecifically during the nESI process, generating artifactual peaks in the mass spectra corresponding to a two ligand-bound complex (Fig. <ref type="figure">5A</ref>). We speculate that in excess concentrations of G10, the molecules can form double helices that are accommodated by the RbCBM74 binding site but that the single molecule binding event represents the most common binding conformation (Fig. <ref type="figure">5B</ref>).</p><p>Because Sas6 encodes both a CBM74 and a CBM26, and this co-occurrence is evolutionarily well-conserved, we speculated that RbCBM26 and RbCBM74 could either bind separate G10 molecules or that one ligand could span the region between the two CBM binding sites <ref type="bibr">[22]</ref>. We used native mass spectrometry to determine the number of G10 molecules bound to Sas6T, which includes both CBMs. The binding state distribution was markedly different when RbCBM26 was included (Fig. <ref type="figure">5C</ref>). At low G10 concentrations, there is a mix of unliganded, 1bound, and 2-bound states unlike BIg-RbCBM74-BIg alone (Fig. <ref type="figure">5D</ref>). As G10 increases, the apo and 1-bound states decrease as the 2-bound fraction increases. For Sas6T, Kd values for 1:1 and 1:2 protein:ligand complexes were calculated to be 3.4 &#177; 0.5 &#181;M and 165.6 &#177; 38.8 &#181;M, respectively, and are in reasonable agreement with ITC data (Extended Data Table <ref type="table">1B</ref>).</p><p>Together these results suggest that RbCBM26 and RbCBM74 each bind one molecule of G10 independently in solution. In the context of a starch granule, this supports a model whereby each CBM of Sas6 binds adjacent &#945;-glucan chains rather than attaching to the same chain in a continuous manner. Moreover, the propensity for BIg-RbCBM74-BIg to bind a single helix of G10 at low ligand concentrations, as also observed with ITC, suggests that this binding platform prefers single helical &#945;-glucans such as amylose, though it can also tolerate double helical stretches of amylopectin. Finally, the sixth cluster (walnut) covers CBM74 domains found in GH13_19 &#945;-amylases. In total, CBM74 domains occur in &#945;-amylases from several subfamilies or non-catalytic dockerincontaining proteins and are widely represented among Bifidobacteria.</p><p>We mapped the conservation of all 99 CBM74 family members onto our structure using CONSURF <ref type="bibr">[54]</ref><ref type="bibr">[55]</ref><ref type="bibr">[56]</ref> (Fig <ref type="figure">6B</ref>). While the central &#946;-sandwich, ion-coordination sphere, and ligand binding site are highly conserved, the flexible loop in RbCBM74 (residues 373-384) occluding the binding site is more variable (Fig. <ref type="figure">6C</ref>; Extended Data Fig. <ref type="figure">9B</ref>). In all but the three or four most closely related CBM74 sequences -covering only the two genera of Ruminococcus and Eubacterium -this loop is short or not present, though how this feature correlates with binding is unknown.</p><p>Most of the key aromatic residues that mediate starch-binding in RbCBM74 are highly conserved (Fig. <ref type="figure">6C</ref>). W373 from RbCBM74 is 100% conserved among all 99 identified CBM74 </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head><p>CBMs are distinct protein domains that assist with substrate breakdown by specifically binding polysaccharide targets. These domains are especially important for binding to insoluble substrates like crystalline cellulose and semi-crystalline starch granules. The CBM74 family binds insoluble starch and its constituents, amylose and amylopectin. CBM74 domains are frequently (81/99 sequences) encoded adjacent to another starch-binding CBM family, either a CBM25 or CBM26 <ref type="bibr">[22]</ref>. Sas6 includes both a CBM26 and a CBM74 domain that have different affinities for maltooligosaccharides but work together to bind granular starch. RbCBM26 has a canonical binding platform that accommodates motifs found in linear and circular maltooligosaccharides. In contrast, RbCBM74 has an extended ligand binding groove that requires at least 8 glucose residues and accommodates the single helices of amylose and the double helices found in amylopectin. Because it is on the cell surface, the CBM74 domain of Sas6 may target R. bromii to the crystalline regions of starch granules that are not easily accessible to human or other bacterial amylases.</p><p>Sas6 is a putative R. bromii amylosome component and likely cooperates with amylases and pullulanases via the interaction of its dockerin domain with a cohesin from a scaffoldin protein <ref type="bibr">[20]</ref>. Because Sas6 is found on the cell surface, it could bind cell anchored scaffoldins Sca2 or Sca5, associate with Sca1/Amy4, or bind the cell surface in a dockerin-independent mechanism <ref type="bibr">[20]</ref>. Breakdown of starch by R. bromii relies on the coordinated effort of approximately 40 distinct proteins, of which Sas6 may play an integral part by specifically targeting the helical regions of starch <ref type="bibr">[20]</ref>.</p><p>Unlike R. bromii, resistant starch-utilizing Bifidobacteria encode CBM74-containing multimodular extracellular amylases <ref type="bibr">[9]</ref>. A recent study looked at the amylases that were differentially encoded between Bifidobacterial strains that could bind and degrade starch granules and those that could not <ref type="bibr">[57]</ref>. Resistant Starch Degrading enzyme 3 (RSD3) was differentially encoded in the resistant starch-binding strains. It contains a CBM74 domain and has high activity on high amylose corn starch. RSD3 has an N-terminal GH13 domain followed by CBM74, CBM26, and CBM25 domains. The CBM74-CBM26 motif is present in RSD3 so the structural and functional insights we have gleaned from Sas6 may suggest how these CBMs structurally assist the enzyme with granular starch hydrolysis.</p><p>Although starch is a polymer composed solely of glucose, there is massive variation in granule structure <ref type="bibr">[7,</ref><ref type="bibr">8]</ref>. This is a function of primary structure (i.e. &#945;1,4 or &#945;1,6 linkages), secondary structure (single or double helices) and tertiary structure (helical packing and amylose content), making granules an exquisitely complex substrate <ref type="bibr">[58]</ref>. This complexity is unlocked by only a few specialized gut bacteria, making granular starch a targeted prebiotic <ref type="bibr">[9,</ref><ref type="bibr">15,</ref><ref type="bibr">16]</ref>. CBM74 domains might serve as a molecular marker for the ability to break down resistant starch in metagenomic samples <ref type="bibr">[22]</ref>. Furthermore, CBM74 domains might make attractive additions to engineered enzymes for enhanced starch degradation on the industrial scale, or as an adjunct to starch prebiotics. The structural and functional picture of RbCBM74 here will accelerate the targeted use of this domain for various health and industrial applications.   2+ atoms are shown as yellow spheres. E. ACX bound at RbCBM26 (green) in chain A (left) and chain B (right), demonstrating minor conformational flexibility that places S286 from RbCBM74 (blue) within the binding site. Side chains involved in ligand binding are shown as green sticks with a hydrogen bond cutoff of 3.2&#197;. ACX is displayed as wheat sticks. Omit map is contoured to 2.0&#963; and carved within 1.6&#197; of ACX ligand. F. RbCBM74 drives binding to granular potato and corn starch. Binding to granular starch was determined by isotherm depletion. The &#956;moles of protein bound per gram of starch was plotted against [free protein] to determine dissociation constants (Kd) and binding maxima (Bmax) using a one-site specific binding model in GraphPad prism. G. Affinity PAGE of Sas6T or individual domains, RbCBM26 and BIg-RbCBM74-BIg, with 0.1% polysaccharide. BSA= bovine serum albumin.  Evolutionary tree for the CBM74 family including 33 sequences selected from the entire studied set of 99 CBM74s (Extended Data Table <ref type="table">2</ref>). Two experimentally characterized CBM74s are marked by an asterisk: Sas6 from Ruminococcus bromii (No. 28, blue cluster) and the subfamily GH13_32 &#945;-amylase from Microbacterium aurum (No. 52; cyan cluster). Protein labels include the order number (33 selected from 1-99), GenBank accession number, abbreviation of the source protein/enzyme and organism name. The tree is based on the alignment (shown in C) spanning the complete CBM74 sequences. B. Structure of RbCBM74 (PDB 7uwv) colored by conservation score from least conserved (green) to most conserved (purple) generated using CONSURF. C. Sequence alignment of the CBM74 family. The six individual groups distinguished from each other by different colors correspond to six clusters seen in the evolutionary tree (panel A); the sequence order in the alignment reflects their order in the tree in the anticlockwise manner (starting from the first sequence in the red cluster). The residues responsible for stacking interactions and involved in hydrogen bonding with glucose moieties of the bound &#945;-glucan are signified by a hashtag and a dollar sign, respectively, above the alignment. The flexible loop observed in the three-dimensional structure of the RbCBM74 is highlighted by the short yellow strip over the alignment. Identical and similar positions are signified by asterisks and dots/semicolons under the alignment blocks. The color code for the selected residues: W, yellow; F, Y -blue; V, L, I -green; D, E -red; R, K -cyan; H -brown; C -magenta; G, P -black. The alignment of all 99 CBM74 sequences of the present study shown in Extended Data Figure <ref type="figure">9B</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>METHODS</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Recombinant Protein Cloning and Expression</head><p>We used a previously described cloning and expression protocol to generate each of the recombinant protein constructs used in this study <ref type="bibr">[59]</ref>. Genomic DNA was isolated from R.</p><p>bromii strain L2-63 and the constructs for Sas6 without the signal peptide were amplified using the primers listed in Table <ref type="table">S1</ref> with overhangs complementary to the Expresso T7 Cloning &amp; Expression System N-His pETite vector (Lucigen). The forward primers were engineered to include the 6x His sequence that complemented the vector plus a TEV protease recognition site for later tag removal. PCR was performed with Flash PHUSION polymerase (ThermoFisher).</p><p>The amplified products and the linearized N-his pETite vector were transformed in HI- </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Sas6 Immunofluorescence</head><p>Custom a-Sas6T antiserum was generated by rabbit immunization with purified recombinant Sas6T protein (Lampire Biological Laboratories). The resulting antiserum was used for western blotting and cell staining. R. bromii cells were grown to mid-log phase on RUM media <ref type="bibr">[17]</ref> with 0.1% potato amylopectin and 2mL of the cell culture was collected for immunostaining and western blotting. For immunostaining, 1mL of R. bromii culture was centrifuged for 1min at 13,000xg and washed 3 times with 1X phosphate buffed saline pH 7.4 (PBS). 2&#181;L of cells were then spread on a glass slide and fixed with 10% formaldehyde in PBS. Slides were washed 3x in PBS to remove fixative but were not permeabilized. Cells were blocked for 30min with 10% goat serum (Jackson ImmunoResearch). a-Sas6T antiserum was diluted 1:1000 in 10% goat serum and applied for 1hr to cells at room temperature. The primary antiserum was removed, and slides were washed 3 x 5min in PBS before the application of 1:500 goat a-rabbit AlexaFluor488 Blocking Buffer (BioRad) for 30min then washed with PBS pH 7.4 + 0.05% Tween 20 (PBST). To detect Sas6, one membrane was incubated with custom rabbit &#945;-Sas6 antiserum (Lampire) diluted 1:500 and the other with custom rabbit &#945;-glutamic acid decarboxylase from R. bromii (Lampire) diluted 1:10,000 in PBST + 5% non-fat dry milk (PBST-milk) for 1hr. Blots were washed in PBST and incubated in horse radish peroxidase-conjugated goat &#945;-rabbit antibody (ThermoFisher) diluted 1:5,000 in PBST-milk and the signal was detected by ECL chemiluminescence (ThermoFisher).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Granular starch binding assays</head><p>Granular starch-binding assays were conducted with potato starch (Bob's Red Mill), corn starch (Sigma), wheat starch (Sigma), or Avicel (Fluka). Prior to use, all polysaccharides were washed 3x with an excess of assay buffer (20mM HEPES pH 7.0, 100mM NaCl) to remove soluble starch and oligosaccharides and prepared as a 50mg/mL slurry. 1mg (corn) or 5mg (potato) of starch slurry was aliquoted into 0.2mL tubes in triplicate, centrifuged at 2,000xg for 2 min and the supernatant was carefully removed. 100&#956;L of protein ranging from 0.5&#956;M-10&#956;M protein was added to each starch and the tubes were agitated by end-over-end rotation at room temperature for 1hr. After centrifugation at 2,000xg for 2min, 20&#956;L of the supernatant was removed for unbound protein concentration determination by absorbance at A280 using a ThermoFisher NanodropOne with three replicate measurements per sample. The remaining 80&#956;L of supernatant was removed and set aside for SDS-PAGE gel analysis. The concentration of unbound protein remaining in the supernatant was used to determine the &#181;moles of protein bound per gram of starch which was plotted against the concentration of initial (free) protein to generate a binding curve <ref type="bibr">[31]</ref>. Overall affinity (Kd) and binding maximum (Bmax) was determined via a one-site binding model (specific binding) using GraphPad Prism version 9.2.0 for Windows (GraphPad Software, San Diego, California USA, www.graphpad.com) <ref type="bibr">[31]</ref>.</p><p>To assess the remaining starch granules for bound protein, the granules were washed three times with an excess of assay buffer by mixing and centrifugation, the final wash supernatant was removed, and 100&#956;L of Laemmli buffer containing 1M urea was added to the starch pellet to denature any bound protein but keep the original volume consistent. To qualitatively determine the amount of unbound and bound protein, 10&#956;L each of the wash supernatant and solubilized pellet fraction were run separately via SDS-PAGE. Bovine serum albumin was used as a negative control and to confirm unbound protein was sufficiently washed from the starch granules.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Polysaccharide Affinity PAGE</head><p>Non-denaturing polyacrylamide gels with and without potato amylopectin (Sigma), corn amylopectin (Sigma), potato amylose (Sigma), bovine liver glycogen (Sigma), pullulan (Sigma),</p><p>or dextran (Sigma) to a final concentration of 0.1% polysaccharide were cast. All polysaccharides were autoclaved and amylose was solubilized by alkaline solubilization with 1M NaOH and acid neutralization to pH 7 with HCl <ref type="bibr">[60]</ref>. Sas6 protein samples were mixed with 6X loading dye lacking SDS. Gels were run concurrently for 4 hours on ice and subsequently stained with Coomassie (0.025% Coomassie blue R350, 10% acetic acid, and 45% methanol). Gels were imaged on a Bio-Rad Gel Doc Go imaging system. The distance between each band and the top of the separating gel were measured using ImageJ <ref type="bibr">[61]</ref>. The ratio of the distance migrated by each band was determined to the distance the BSA band traveled. Binding was considered positive if the ratio was less 0.85 as previously described <ref type="bibr">[62]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Isothermal Titration Calorimetry</head><p>All ITC experiments were carried out using a TA Instruments standard volume NanoITC. For each experiment, 1300&#956;L of 25&#956;M protein was added to the sample cell and the reference cell was filled with distilled water. The sample injection syringe was loaded with 250&#956;L of the appropriate ligand concentration (0.5mM -5mM) to fully saturate the protein by the end of 25 injections of 10&#181;ls. Titrations were performed at 25&#176;C with a stirring speed of 250 rpm. The resulting data were modeled using TA Instruments NanoAnalyze software employing the pre-set models for independent binding and blank (constant) to subtract the heat of dilution. For interactions with high affinity (c-value at 25&#181;M protein greater than 5), no alterations were made to the model. If the calculated c value of an interaction fell below 5, the n value was set to 1 as indicated in the figure legend following the guidance for modeling low affinity interactions <ref type="bibr">[63]</ref>. For polysaccharide titrations, curves were modeled by varying the substrate concentration until n=1 such that the Kd represents the overall affinity for the construct <ref type="bibr">[52]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Protein Crystallization</head><p>Crystallization conditions for &#945;-cyclodextrin (2mM) bound (pdb 7UWW) and unliganded (pdb 7UWU) crystals of Sas6T were screened via 96-well sparse matrix screen (Peg Ion HT, Hampton</p><p>Research #HR2-139) in a sitting drop vapor diffusion experiment at room temperature. Screens were set up using an Art Robbins Gryphon robot with 20mg/mL protein in a 3-well tray (Art Robbins #102-0001-13) using protein-to-well solution ratios of 2:1, 1:1, and 1: </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structure Determination and Refinement</head><p>X-ray data were collected at the Life Sciences Collaborative Access Team (LS-CAT) at Argonne National Laboratory's Advanced Photon Source (APS) in Argonne, IL. Data were processed at APS using autoPROC with XDS for spot finding, indexing, and integration followed by Aimless for scaling and merging <ref type="bibr">[64]</ref><ref type="bibr">[65]</ref><ref type="bibr">[66]</ref>. Intrinsic sulfur SAD phasing was used to determine the structure of Sas6T/&#945;-cyclodextrin (7UWW) using AutoSol in Phenix <ref type="bibr">[67,</ref><ref type="bibr">68]</ref>. Those coordinates were then used for molecular replacement in Phaser to determine the unliganded Sas6T (7UWU) and BIg-RbCBM74-BIg/G10 (7UWV) structures <ref type="bibr">[69]</ref>. All three structures were refined via manual model building in Coot and refinement in Phenix.refine <ref type="bibr">[70,</ref><ref type="bibr">71]</ref>. Metal ion identities were validated using the web-based CheckMyMetal (CMM) tool <ref type="bibr">[72]</ref> (<ref type="url">https://cmm.minorlab.org/</ref>). Carbohydrate models were validated using Privateer <ref type="bibr">[73]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>SEC-SAXS experiment</head><p>SAXS was performed at Biophysics Collaborative Access Team (BioCAT, beamline 18ID at APS) with in-line size exclusion chromatography (SEC-SAXS) to separate the sample from aggregates and other contaminants. Sample was loaded onto a Superdex 200 Increase 10/300 GL column (Cytiva), which was run at 0.6ml/min by an AKTA Pure FPLC (GE) and the eluate after it passed through the UV monitor was flown through the SAXS flow cell. The flow cell consists of a 1.0mm</p><p>ID quartz capillary with ~20&#956;m walls. A coflowing buffer sheath is used to separate the sample from the capillary walls, helping prevent radiation damage <ref type="bibr">[74]</ref>. Scattering intensity was recorded using a Pilatus3 X 1M (Dectris) detector which was placed 3.6m from the sample giving a q-range of 0.003&#197; -1 to 0.35&#197; -1 . 0.7 s exposures were acquired every 1s during elution and data was reduced using BioXTAS RAW 2.1.1 <ref type="bibr">[75]</ref>. Within RAW, the Volume of Correlation (VC), molecular weight, and oligomeric state were determined <ref type="bibr">[76,</ref><ref type="bibr">77]</ref>. Buffer blanks were created by averaging regions flanking the elution peak and subtracted from exposures selected from the elution peak to create the I(q) vs q curves used for subsequent analyses. The molecular weight was calculated by comparison to known structures (Shape&amp;Size) <ref type="bibr">[38]</ref>. P(r) function was determined using GNOM <ref type="bibr">[39]</ref>. GNOM and Shape&amp;Size are part of the ATSAS package (version 3.0) <ref type="bibr">[78]</ref>. High resolution structures were fit to the SAXS data using FoXS and flexibility in the high-resolution structures was modeled against the Multi-FoXS data <ref type="bibr">[40]</ref>. Tables S2A-C list sample, instrumentation, and software for the SEC-SAXS experiment.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Hydrogen-Deuterium eXchange Mass Spectrometry (HDX-MS)</head><p>HDX-MS experiments were performed using a Synapt G2-SX HDMS system (Waters), similar to previously reported <ref type="bibr">[79]</ref>. Deuteration reactions were incubated at 20&#176;C for 15s, 150s, 1500s, and 15,000s in triplicate. 3&#956;L of BIg-RbCBM74-BIg alone or in the presence of G10 were diluted with 57&#956;L of deuterated labeling buffer. Nondeuterated data were acquired by dilution with protonated buffer and fully deuterated data were prepared by dilution in 99% D2O, 1% (v/v) formic acid) for 48h at room temperature. Samples were measured in triplicate using automated handling with a PAL liquid handling system (LEAP), using randomized sequential collection with Chronos.</p><p>Following incubation, deuteration was quenched by mixing 50&#956;L of the solution with 50&#956;L of 100mM phosphate, pH 2.5 at 0.3&#176;C. Immediately after the samples were quenched, 95&#956;L of the sample was loaded onto an Acquity M-class UPLC (Waters) with sequential inline pepsin digestion (Waters Enzymate BEH Pepsin column, 2.1mm &#215; 30mm) for 3min at 15&#176;C followed by reverse phase purification (Acquity UPLC BEH C18 1.7&#956;m at 0.2&#176;C). Sample was loaded onto the column equilibrated with 95% water, 5% acetonitrile, and 0.1% formic acid at a flow rate of 40&#956;L/min. A 7min linear gradient (5%-35% acetonitrile) followed by a ramp and 2min block (85% acetonitrile) was used for separation and directly continuously infused onto a Synapt XS using Ion Mobility (Waters). [Glu1]-Fibrinopeptide B was used as a reference.</p><p>and the S-lens RF level was kept at 80. Low m/z detector optimization and high m/z transfer optics were used, and the trapping gas pressure was set to 2. In-source trapping was enabled with the desolvation voltage fixed at -25V for improved ion transmission and efficient salt adduct removal.</p><p>Transient times were set at 128ms (resolution of 25,000 at m/z 400), and 5 microscans were combined into a single scan. A total of ~50 scans were averaged to produce the presented mass spectra. All full scan data were acquired using a noise threshold of 0 to avoid pre-processing of mass spectra. A total of three measurements for each ligand concentration were performed. Data were then processed and deconvoluted using UniDec software <ref type="bibr">[81]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Kd Measurements by Native MS.</head><p>We performed titration experiments for both BIg-RbCBM74-BIg and Sas6T using G10 and acquired modeled titration curves. Each bound state differed by ~1639 Da, which agrees with the theoretical mass of G10. To obtain the binding constants, we summed the peak intensities of all abundant charge states in our mass spectra. Kd values were calculated using the relative intensities of unbound protein and each ligand bound species from the mass spectra as previously described <ref type="bibr">[82]</ref>. Briefly, the protein-ligand binding equilibrium of BIg-RbCBM74-BIg with G10 in solution can be described by the following reversible reaction:</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>"( "%(</head><p>where ! is the ligand and " and "! are the free protein and protein with one specifically bound ligand, respectively. BIg-RbCBM74-BIg possesses one ligand-binding site, RbCBM74. As the concentration of ligand is increased, ligand molecules can bind nonspecifically during the nESI process, generating artifactual peaks in the mass spectra corresponding to a two ligand-bound complex. As the concentration of ligand is increased, ligand molecules can bind nonspecifically the titration experiment at each ligand concentration and can then be related to the equilibrium constants:</p><p>[L] can also be determined from nESI-MS titration data:</p><p>[L] was then obtained at each ligand concentration and applied to the Eqs. 4a-c. Equations 4a-b were then fitted to experimental fractional intensities using nonlinear least-squares curve fitting using the lsqnonlin.m. function in MATLAB. A more detailed derivation of these equations is provided elsewhere <ref type="bibr">[82]</ref>, along with the approach utilized for Sas6 which possesses two sites for specific binding (RbCBM74 and RbCBM26) and exhibits a third nonspecific bound state as shown in Eq. 6.</p><p>! ! " &#8652; "% &#8652; "% 7 &#8639;&#8642; ! &#8639;&#8642; ! &#8639;&#8642; !</p><p>"( &#8652; "%( &#8652; "% 7 (</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Sequence collection</head><p>Amino acid sequences of CBM74 modules were collected according to information in the CAZy </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Sequence comparison and evolutionary analysis</head><p>The alignment of 99 CBM74 modules from the final set was performed using the program Clustal-Omega (<ref type="url">https://www.ebi.ac.uk/Tools/msa/clustalo/</ref>) <ref type="bibr">[90]</ref>. Only a subtle manual tuning of the computer-produced alignment was necessary to perform to maximize sequence similarities. The evolutionary tree of these 99 sequences was calculated by a maximum-likelihood method (on the final alignment including the gaps) using the WAG substitution model and the bootstrapping procedure with 500 bootstrap trials implemented in the MEGA-X package <ref type="bibr">[91]</ref><ref type="bibr">[92]</ref><ref type="bibr">[93]</ref>. The calculated tree file was displayed with the program iTOL <ref type="bibr">[94]</ref> (<ref type="url">https://itol.embl.de/</ref>). From both the alignment and the tree of all 99 sequences, a sample of 33 representative CBM74s was selected for a simplified alignment and tree. The structural comparison was created using the above-mentioned alignment in conjunction with the web-based CONSURF tool <ref type="bibr">[54]</ref><ref type="bibr">[55]</ref><ref type="bibr">[56]</ref>. Extended Data Figure <ref type="figure">9</ref>: Conservation of binding residues among all 99 CBM74 family members. A. A maximum-likelihood tree covering 99 sequences with emphasis on the two experimentally characterized CBM74s, Sas6 from Ruminococcus bromii (No. 28, blue cluster) and the subfamily GH13_32 &#945;-amylase from Microbacterium aurum (No. 52; cyan cluster). For details concerning all 99 CBM74 sequences, see Extended Data Table <ref type="table">2</ref>. A simplified tree showing 33 selected CBM74 sequences representing all clusters is shown in Fig. <ref type="figure">6A</ref>. B. Sequence alignment of the 99 CBM74 sequences. The labels of protein sources consist of the order number (1-99), GenBank accession number, abbreviation of the source protein/enzyme and the name of the organism. The two experimentally characterized CBM74 are marked by an asterisk. The six individual groups distinguished from each other by different colors correspond to six clusters seen in the evolutionary tree (panel A); the sequence order in the alignment (starting from the top from 1 to 99) reflects their order in the tree in the anticlockwise manner (starting from the first sequence in the red cluster). The residues responsible for stacking interactions and involved in hydrogen bonding with glucose moieties of the bound &#945;-glucan are signified by a hashtag and a dollar sign, respectively, above the alignment. The flexible loop observed in the three-dimensional structure of RbCBM74 is highlighted by the short yellow strip over the alignment. Identical and similar positions are signified by asterisks and dots/semicolons under the alignment blocks. The color code for the selected residues: W, yellow; F, Y -blue; V, L, I -green; D, E -red; R, K -cyan; H -brown; C -magenta; G, P -black.</p></div></body>
		</text>
</TEI>
