<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Multiple Choice vs. Fill-In Problems: The Trade-off Between Scalability and Learning</title></titleStmt>
			<publicationStmt>
				<publisher>ACM</publisher>
				<date>03/18/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10541413</idno>
					<idno type="doi">10.1145/3636555.3636908</idno>
					
					<author>Ashish Gurung</author><author>Kirk Vanacore</author><author>Andrew A Mcreynolds</author><author>Korinn S Ostrow</author><author>Eamon Worden</author><author>Adam C Sales</author><author>Neil T Heffernan</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Learning experience designers consistently balance the trade-off between open and close-ended activities. The growth and scalability of Computer Based Learning Platforms (CBLPs) have only magnified the importance of these design trade-offs. CBLPs often utilize close-ended activities (i.e. Multiple-Choice Questions [MCQs]) due to feasibility constraints associated with the use of open-ended activities. MCQs offer certain affordances, such as immediate grading and the use of distractors, setting them apart from open-ended activities. Our current study examines the effectiveness of Fill-In problems as an alternative to MCQs for middle school mathematics.We report on a randomized study conducted from 2017 to 2022, with a total of 6,768 students from middle schools across the US. We observe that, on average, Fill-In problems lead to better post-test performance than MCQs; albeit deeper explorations indicate differences between the two design paradigms to be more nuanced. We find evidence that students with higher math knowledge benefit more from Fill-In problems than those with lower math knowledge.
CCS CONCEPTS• Applied computing → Computer-assisted instruction; Interactive learning environments.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">INTRODUCTION</head><p>The rapid growth in technology and the ability to produce educational material accessible by innumerable learners has led to the development and adoption of Computer Based Learning Platforms (CBLPs) across educational sectors. With access to the internet, learners across the world use CBLPs in the form of Massive Open Online Courses (MOOCs), Learning Management Systems (LMSs), and standalone online learning platforms. The past two decades have seen a drastic rise in the implementation and utilization of online educational materials <ref type="bibr">[17,</ref><ref type="bibr">23]</ref>. With this growth, Learning Experience (LX) designers are tasked with the critical responsibility of ensuring that learning materials are not only effective but also scalable. A key difficulty faced by LX designers is finding a balance between the educational merits of different problem types and the feasibility of employing these problem types effectively at scale. This balance is crucial for maximizing the impact and reach of these CBLPs.</p><p>Broadly, the instructional strategies leveraged by LX designers in terms of the problem types can be classified into two categories: open-ended and closed-ended problems. Closed-ended problems, such as Multiple Choice Questions (MCQs), 'Check all that Apply', and 'Arrange in the Correct Order', lend themselves more easily to the integration of automated grading and instant feedback. These features, along with on-demand help, can significantly enhance the learning experience for users. This approach allows for a more scalable and efficient way of delivering instruction, particularly in educational settings, with automation playing a key role in reducing the demand on instructors' time and resources <ref type="bibr">[4]</ref>. Alternatively, incorporating automated grading and instant feedback in openended problems such as short answer questions, essays, and fill-in problems is more challenging. While closed-ended problems are more suitable for automation compared to open-ended problems, researchers have raised concerns regarding their use <ref type="bibr">[12,</ref><ref type="bibr">32,</ref><ref type="bibr">49]</ref>. They point out that closed-ended problems can be susceptible to recognition and guessing, which can lead to shallow learning. In contrast, open-ended problems are often viewed as more rigorous and allow the instructor to infer the learners' understanding and comprehension of the topic from their answers <ref type="bibr">[12,</ref><ref type="bibr">49]</ref>. While rigor and thoroughness are highly desirable in educational settings, it is also important to acknowledge that the higher level of difficulty and the demand for rigorous engagement can be cognitively taxing on the learners. This strain often stems from the need to exercise recall over recognition, and generation over selection, which inherently requires higher cognitive effort.</p><p>Research into the comparative effectiveness of close-ended and open-ended activities in facilitating learning is somewhat mixed. Some studies have found a preference for traditional open response problems (ORP) over MCQs <ref type="bibr">[1,</ref><ref type="bibr">12,</ref><ref type="bibr">32]</ref>, while others underscore the merits of MCQs <ref type="bibr">[16,</ref><ref type="bibr">48,</ref><ref type="bibr">50]</ref>. Furthermore, others suggest there's little to no difference between the two formats <ref type="bibr">[30,</ref><ref type="bibr">31,</ref><ref type="bibr">44,</ref><ref type="bibr">49]</ref>. Amidst this backdrop of varied and sometimes conflicting findings when comparing traditional ORPs to MCQs, Fill-In problems emerge as an intriguing point of discussion. Fill-In problems share several characteristics with MCQs, including the benefits of automated grading, immediate feedback, and availability of on-demand help. Furthermore, a distinct and linear relationship exists between MCQs and Fill-In problems, enabling the straightforward conversion of one format to the other, thus offering versatility in the design of learning activities.</p><p>In this paper, we investigate the application of Fill-In problems as an alternative to MCQs in mastery-based activities. To this end, we conducted an in-vivo randomized study aimed at exploring the relative difficulty of utilizing Fill-In problems as an alternative to MCQs. We then explore the influence of these problem types on learners' performance in mastery-based activities and a posttest with a more complex transfer task upon acquiring mastery. Finally, we explore the potential heterogeneity in the effectiveness of the problem types across learners with varying mathematical prior performances. Specifically, we explore the following research questions:</p><p>(1) Is there a difference in the difficulty between equivalent MCQ and Fill-In problems? (2) Does the use of different problem types (MCQ vs. Fill-In) on mastery-based activities impact students' learning? (3) How does the effectiveness of problem type vary among learners with differing levels of mathematical proficiency?</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">PRIOR WORKS 2.1 MCQs &amp; Fill-In Problems</head><p>Over the years, various prior research has explored the efficacy of utilizing MCQs over ORPs and found mixed results, with some finding MCQs to be more beneficial <ref type="bibr">[16,</ref><ref type="bibr">48,</ref><ref type="bibr">50]</ref>, others finding ORPs to be more beneficial <ref type="bibr">[1,</ref><ref type="bibr">12,</ref><ref type="bibr">32]</ref>, and others finding little to no difference between the two problem types <ref type="bibr">[30,</ref><ref type="bibr">31,</ref><ref type="bibr">44,</ref><ref type="bibr">49]</ref>. However, prior exploration regarding the feasibility of the two problem types has shown ORPs to be more costly towards instructor resources than MCQs <ref type="bibr">[4]</ref>. Beyond their usage, it is also crucial to acknowledge that other contextual factors can influence the use of one problem type over the other, as each has its unique advantages and disadvantages. For example, MCQs can be particularly beneficial in assessing large cohorts of students en masse( i.e. SAT, TOEFL) <ref type="bibr">[26,</ref><ref type="bibr">35,</ref><ref type="bibr">38]</ref>. On the other hand, ORPs have been shown to provide a more accurate assessment of learners' understanding and comprehension in various STEM-related subjects compared to MCQs <ref type="bibr">[16,</ref><ref type="bibr">48]</ref>. In fact, prior studies have reported on inflation of grades when utilizing MCQs over ORPs in STEM-related subjects <ref type="bibr">[16,</ref><ref type="bibr">48]</ref>.</p><p>While MCQs are widely used in CBLPs, researchers have expressed concerns regarding their usage. MCQs can be susceptible to synthesis, guessing, and recognition due to the use of distractors<ref type="foot">foot_0</ref> resulting in shallow learning <ref type="bibr">[10,</ref><ref type="bibr">19,</ref><ref type="bibr">35]</ref>. Consequently, there have been concerns regarding the reliability and validity of the use of MCQs when inferring learners' knowledge and ability <ref type="bibr">[14]</ref>. In particular, the presence of distractors in MCQs can inadvertently trigger learners' recall of topics they are attempting to work on, thereby diminishing the effectiveness of MCQs in fostering more rigorous learning and developing a deeper understanding of the topic <ref type="bibr">[42]</ref>. Furthermore, MCQs, due to their design, are relatively less conducive to fostering creative thinking and idea generation <ref type="bibr">[9]</ref>, which are crucial skills for comprehensive learning. In contrast, Open Response Problems (ORPs) inherently require students to demonstrate higher-level thinking and reasoning for each problem, thereby eliminating the guessing element commonly associated with MCQs <ref type="bibr">[35]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Mastery-Based Learning Activities</head><p>In recent decades mastery-base learning activities have become a common pedagogical technique, especially in CBLPs. Rather than assuming learning upon completion of certain activities associated with the material, mastery-based learning requires learners to demonstrate knowledge and understanding of the concepts before progressing to the next topic <ref type="bibr">[8]</ref>. Mastery-based learning approaches have shown to reduce variance in student aptitude <ref type="bibr">[2,</ref><ref type="bibr">29]</ref>, increase long-term retention of knowledge <ref type="bibr">[29]</ref>, change student attitude towards content <ref type="bibr">[2,</ref><ref type="bibr">29]</ref>, and increase self-belief <ref type="bibr">[2,</ref><ref type="bibr">18]</ref>.</p><p>One of the primary features of mastery-based learning is to provide students with the ability to practice the skills that allows the teachers to assess their students' abilities while facilitating learning opportunities. CBLPs, by design, have an advantage when implementing mastery-based assignments, as the activity can adapt to the student's performance. Various CBLPs have explored the implementation of mastery-based assignments using various approaches. While some platforms, such as Khan Academy <ref type="bibr">[28,</ref><ref type="bibr">34]</ref>, and ASSISTments <ref type="bibr">[20]</ref>, have explored using an arbitrary threshold of N-Consecutive Correct Responses (N-CCR), others have relied on more precise measures of mastery using Knowledge Tracing (KT) models <ref type="bibr">[13]</ref>. KT models predict student performance in future problems by leveraging their past performance on similar or related skills. Both N-CCR and KT approaches have their merits and flaws; N-CCR is more explainable and interpretable by teachers, whereas KT models are harder to understand for the teachers but are more accurate at estimating learner mastery. While a heuristic of N-CCR could be considered rather simplistic, <ref type="bibr">Kelly et al. (2015)</ref>  <ref type="bibr">[25]</ref> reported that a N-CCR design, with N = 3, has comparable performance in estimating mastery to more sophisticated KT models. Additionally, <ref type="bibr">Prihar et al. (2022)</ref>  <ref type="bibr">[36]</ref> have reported on the benefits of using N = 2, 4, and 5 as a threshold and found N = 3 to be an optimal threshold. While a simple N-CCR design is easy to implement <ref type="bibr">[22,</ref><ref type="bibr">25]</ref> and interpret, some have expressed concerns regarding the risks of inequitable outcomes due to the use of N-CCR design's assumption across students with different learning rates <ref type="bibr">[15]</ref>. Although concerns about the use of N-CCR and its potential impact on creating inequitable outcomes are significant, the study by Koeginder et al. (2023) <ref type="bibr">[27]</ref>, which reports a surprising consistency in students' learning rates under ideal conditions suggests a possible avenue to both mitigate the concerns related to inequity stemming from varying learning rates and an opportunity to revisit the risk of inequity in outcomes on mastery-based activities due to the methodology utilized in estimating mastery.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">CURRENT STUDY 3.1 Experimental Design</head><p>The current study comparing MCQs and Fill-In problems was conducted using ASSISTments <ref type="bibr">[20]</ref>, a CBLP popular among middle school math teachers in the United States. In this experiment, we developed two mastery-based activities focused on the mathematical concepts of 'Greatest Common Factor' (GCF) and 'Evaluating Expressions' (EE). These activities were designed in accordance with the Common Core State Standards <ref type="bibr">[33]</ref>, with the GCF activity developed using the grade 6 curriculum and the EE activity developed using the grade 7 curriculum.</p><p>As illustrated in Figure <ref type="figure">1</ref>, each problem set in our study included a mastery learning component followed by a post-test. The students are randomized to one of two problem types in the mastery learning components: MCQs or Fill-In problems. A N-CCR design with N=3 for the mastery-based activity is utilized in both conditions to estimate mastery, i.e., students need to correctly answer 3 problems in a row to demonstrate their mastery of the content. If a student is incorrect on their first attempt or asks for hints, the consecutive correctness counter is reset to zero. During the assignment, students have the option to request up to three hints, with the bottomout hint giving away the answer to the problem. Additionally, the system also imposes a daily limit of 10 problems per condition. However, if a student correctly answers the 9th or 10th problem, they are allowed to attempt up to 11 or 12 problems, respectively, to demonstrate mastery. Students unable to demonstrate mastery within the first 10 problems are required to wait until the following day to continue with the activity.</p><p>Upon demonstrating mastery, students are asked to take a twoproblem post-test. These problems are transfer items on the same topic as the experiment. These items required that students to apply the mastered knowledge component to a relatively more complex problem on the post-test. Examples of the problems in the experiment (MCQs vs. Fill-In problems) and the post-test are illustrated in Figure <ref type="figure">2</ref>. As demonstrated in Figure <ref type="figure">2</ref>, post-test problems are relatively more complex in comparison to the problems in the mastery learning component. The problem complexity was increased on the post-test by increasing the dimensionality of the problem from 2 to 3 and using a more complex sentence structure on the problem. The objective here is to assess the performance of the students who demonstrated mastery on a more complex transfer task as a representation of their learning <ref type="bibr">[47]</ref>. We chose to utilize Fill-In questions as they align more closely with the problem type that would be utilized in a traditional experimental setup where the post-test would likely be conducted using a traditional paper and pencil approach to assess the students' learning on the transfer item.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Description of Dataset</head><p>The data was collected across five school years in the United States <ref type="bibr">(2017-18, 2018-19, 2019-20, 2020-21, 2021-22)</ref>. During this time, the assignments were made available to middle school teachers who use ASSISTments as an instructional tool by assigning mastery-based activities to their students as part of their lessons. During our study, 192 teachers assigned the two problem sets to 383 classes. A total of 6774 students participated in the experiment. A small number of students, 20, worked on both problem sets. In such instances, we only included the student data from the first participation and dropped the other records.</p><p>In addition to data on student performance, hint usage, and time to first attempt per problem in the mastery-based components of the experiment, the student performance on the post-test items and the average prior percent correctness across all problems the students worked on the CBLP prior to participating in the experiment (prior performance) were also calculated.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3">Descriptive Statistics</head><p>The descriptive statistics on student behavior in the mastery learning component are presented in Table <ref type="table">1</ref>. While there was no significant difference in the average number of problems taken to reach mastery between the MCQ and Fill-In conditions, other behavioral differences were notable. Specifically, students in the Fill-In condition, on average, accessed more hints and took more time before submitting their first responses compared to their counterparts in the MCQ condition. Intriguingly, despite every incorrect attempt and hint request resulting in a loss of 33% partial credit, students in the Fill-In group achieved a higher average score on the mastery components. This suggests that while students in the MCQ condition likely made more attempts than those in the Fill-In group, the students in the Fill-In condition were able to recognize they needed help, request it, and effectively utilize it to the problem.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4">Analysis Plan</head><p>To address our research questions, we conducted three analyses of the experimental data. The first (Analysis 1: Section 4) addresses the differences in problem difficulty caused by problem type within the mastery learning activity and explores students' performance patterns within each activity. Next (Analysis 2: Section 5), estimates  the effects of the problem type on students learning, as represented by their performance on the post-test. Finally (Analysis 3: Section 6), evaluates whether the effect of problem type varies based on </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">ANALYSIS 1: ASSESSING THE DIFFERENCES IN DIFFICULTY BETWEEN MCQS AND FILL-IN PROBLEMS</head><p>The students were randomized into Fill-In and MCQ conditions with equivalent problems across conditions in the mastery component, i.e., the MCQ and Fill-In problems had the same problem body and answers. The difficulty of the problem type can be evaluated by comparing student performance across conditions. This analysis allows us to understand whether different problem types can influence student performance during mastery-based activities. Furthermore, estimating differences in difficulty can help contextualize potential differences in performance and learning outcomes across conditions.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1">Methods</head><p>The relative difficulty between Fill-In and MCQs was estimated using linear regression models with robust standard errors using the estimator <ref type="bibr">[39]</ref> package in R <ref type="bibr">[37]</ref>. First, we ran a model only utilizing student data from the first problem they attempted in the mastery learning component. This method allows us to isolate the difference in performance caused by problem types from any potential learning or attrition that could occur as the students work through the mastery component. Equation ( <ref type="formula">1</ref>) represents the difficulty estimation model. Let &#119884; &#119894; &#119895; indicate whether student &#119894; was correct on the first attempt of problem &#119895;. Let Fill-In &#119894; indicate whether student &#119894; was randomized to receive Fill-In problems, and &#119875; be an indicator for each problem the students attempted. Since assignment to problems and problem type were randomized and the problems &#119895; were equivalent across conditions, the effect of the Fill-In condition (&#120573; 1 ) is an unbiased causal effect. Note that if a student saw a problem but did not submit a response, we considered their responses to be incorrect. This step ensured that differences in dropout rates did not bias the results.</p><p>To examine how students performed on subsequent problems in the mastery learning component, we reran this analysis for the first 10 problems the students can attempt before reaching the daily limit without exhibiting mastery. Although these are no longer unbiased estimates of causal effects of problem types-due to the potential spillover effects from learning on previous problems and the differences in samples due to acquisition of mastery and attrition rates across conditions-the differences are still informative of students' learning experiences.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2">Results</head><p>Students performance on the first problem in the mastery learning component of the experiment differed significantly based on their treatment assignment such that students in the MCQ condition outperformed those in the Fill-In condition by an estimated five percentage points (&#120573; 1 = -0.05, SE = 0.012, p &lt; 0.0001). This coefficient is an unbiased estimate of the difference in difficulty caused by problem type.</p><p>Figure <ref type="figure">3</ref> displays the average performance (lines) and samples (shading) by condition across the first ten problems in the mastery learning component of the experiment. Notably, for the first problem, students in the MCQ condition perform better than their peers in the Fill-In condition. However, this difference dissipates for the subsequent two problems before reversing, with the students in the Fill-In condition outperforming their peers in the MCQ condition. This finding suggests that the difficulty experienced early on by the students in the Fill-In condition potentially benefited the students later in the activity. Nevertheless, it is important to acknowledge a potential confound to this explanation as the samples may differ across conditions on problem sequences greater than one. This possibility of potential confound is illustrated in Figure <ref type="figure">3</ref>, using the difference in the shaded area (e.g., the percent of students within each condition who started the problem in that problem sequence) on the second and third problem before the students could master the knowledge component. Further exploration of this potential confound is reported in Section 5.2.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">ANALYSIS 2: IMPACT OF FILL-IN PROBLEMS AND MCQS ON STUDENT LEARNING</head><p>Analysis 1 (Section 4) shows that Fill-In problems are more difficult than MCQs. This finding suggests that students tackling Fill-In problems exerted more effort to achieve mastery. This heightened challenge might obscure their recognition of substantial learning progress and potentially lead to negative emotions. Such feelings could result in unproductive behaviors, including gaming <ref type="bibr">[3]</ref>, wheelspinning <ref type="bibr">[6]</ref>, or even dropping out of the activity entirely. Thus, there is the potential that the increased difficulty caused by Fill-In responses may have an adverse effect on student learning. However, Analysis 1 also showed that students in the Fill-In condition outperformed those in the MCQ condition later in the masterly learning activity. Conversely, MCQs may induce students to engage in shallow learning, such as employing educated guesses and deducing answers, effectively recognizing and synthesizing the information. They could potentially gain only perfunctory mastery of the concept by deliberating over the provided choices in the MCQs instead of taking full advantage of the learning opportunities. Thus, it is unclear which problem type is more conducive to learning based on differences in problem difficulty alone.</p><p>In Analysis 2, we evaluate the impact of problem type (MCQ vs Fill-In) within mastery-based activities on student learning. To measure the effectiveness of each approach, we analyze students' performance on the post-test problems. As explained in Section 3.1, these post-tests are designed to assess how well students can apply their recently mastered knowledge to more complex problems within the same topic (i.e. transfer problems).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1">Methods</head><p>Evaluating the impact of problem type on students' learning requires two steps. First, since only the students who mastered the knowledge component took the post-test, and some students who mastered the knowledge did not start or complete the post-test, we must ensure that this attrition does not bias our outcomes. This assessment is critical given that Fill-In problems, being more difficult than their MCQ counterparts, might prompt students with lower knowledge levels to attrit. Once we establish whether attrition is balanced across conditions, we can estimate the effects of the problem types on student learning as measured by their performance on the post-test.</p><p>To evaluate whether attrition was balanced across conditions, we employed two tests. First, we compared the difference in attrition rates between conditions to thresholds established by the U.S. Institute of Education Sciences (IES) <ref type="bibr">[21]</ref>. Next, we conducted a Chi-Square test to estimate whether the attrition differed significantly between conditions.</p><p>To evaluate whether students are more likely to learn while working on Fill-In problems than MCQs, we estimated a mixedeffect logistic regression model using the glmer package in R <ref type="bibr">[5]</ref>. We regress indicators of whether the students got each individual post-test problem correct on the first attempt on a binary indicator for the Fill-In condition. We use this method because averaging the post-test problem together would not have created a continuous variable, as there were only two post-test problems for each problem set. Therefore, a linear regression likely has a poor model fit. Using a logistic regression model, we can treat each post-test problem individually -this also allows us to include students even if they did not complete both post-test problems. We include random interprets for post-test problems to account for differences in problem difficulty and students because students completed multiple post-test problems. We also include random intercepts for the students' classes because their classroom context could influence their learning behaviors, and students are often grouped within classes with students of similar abilities.</p><p>For any given post-test question &#119895; completed by student &#119894;, the model for the likelihood of correctness is represented by Equation <ref type="bibr">(2)</ref> where &#120574; 0 is the fixed intercept, &#120583; &#119894; is the random intercept for each student, &#120583; &#119888; is the random intercept for each student's class, and &#120583; &#119895; is the random intercepts for each post-test problem. Let Fill-In &#119894; be a binary indicator for whether a student is in the Fill-In condition. The coefficient, &#120574; 1 , for the Fill-In problems, represents the difference in likelihood of the students in the Fill-In condition answering the post-test items correctly in comparison to the students in the MCQs condition. Because assignment to condition was random, &#120574; 1 is the causal effect of Fill-In problems on mastery.</p><p>(2)</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2">Attrition</head><p>Table <ref type="table">2</ref> details the experiment's attrition rates and the balance test statistics. The overall attrition rate was 27.07%. Attrition occurred at three distinct levels: Firstly, 18.02% of students did not demonstrate mastery in the learning component and thus were unable to take the post-test. Secondly, 9.91% of students achieved mastery but did not commence the post-test. Thirdly, a subset of students only completed one problem of the post-test. These students were not excluded from our analysis and are not reflected in the overall attrition figure. Based on the results of both the IES threshold and the Chi-squared test, suggesting that attrition was balanced across conditions at all levels. However the differences in mastery rates between the conditions were only marginally non-significant (p = 0.006), thus we choice to incorporate a robustness check of our effects estimation presented below.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.3">Results</head><p>Table <ref type="table">3</ref> presents the parameters of the model used to estimate the effect of problem type on post-test performance. There was a positive causal effect of Fill-In problems on student performance on transfer tasks. Specifically, students who engaged with Fill-In problem sets were significantly more likely to provide correct responses in the post-test compared to those who worked through MCQs (&#120574; 1 = 0.23, SE = 0.06, p &gt; 0.001). Putting this on the probability scale, students in the MCQ condition had a 27% probability of getting either of the transfer problems correct, whereas students in the Fill-In condition had a 31% probability of getting the transfer item correct. Notably, this higher likelihood of correctly answering post-test problems persisted even after adjusting for the variance attributable to individual students, their respective classes, and the problems themselves.</p><p>In our analysis, we also examined the variances of &#120583;, &#120591;, to evaluate the variance in post-test performance attributed to students and their classes. Over one-third, 34% of the variance was associated with the random intercepts. A substantial portion of the performance variance was associated with individual students (&#120591; &#119894; = .83; 17% of the variance ). However, the class environment also played a substantial role, accounting for a considerable proportion of the variance (&#120591; &#119888; = .60; 12% of the variance). This finding indicates the importance of the student's learning environment and peer group in their performance. As part of our analysis of the potential heterogeneity in outcome across different prior knowledge among students, we also delve deeper into the impact of the learning environment on student outcomes in the upcoming section, Section 6.</p><p>Notably, this analysis included only post-test problems that students attempted, thus excluding students who did not master the knowledge component and those who did not start the post-test. Furthermore, some students only completed one problem on the post-test. Although we showed that the differences in attrition across conditions were not significant, it is still possible that they biased our outcomes. Thus, we reran the analysis, coding all of the students who did not master or attrited in any way with a zero for each post-test problem. The results of this robustness check model revealed no difference in significance or magnitude of the causal effect.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6">ANALYSIS 3: HETEROGENEITY IN THE IMPACT OF PROBLEM TYPES DIFFERENT LEVELS OF PRIOR PERFORMANCE</head><p>Finally, to address our last research question of whether the effect of problem type problems on learning varies based on students' prior ability, we ran one final analysis. This allows us to assess a potential nuance of how Fill-In problems benefit student learning. As detailed in Section 4, Fill-In problems were more difficult than the equivalent MCQs. As such, students of varying mathematical proficiency may derive different benefits from each problem type. It's plausible that higher-knowledge students potentially benefit more from the critical thinking, retrieval, and recall required to solve the Fill-In problems, which lack the options provided in MCQs. On the other hand, lower-knowledge students might find the availability of the options in MCQs beneficial in developing intuition and learning the concept, as distractors can highlight potential misconceptions and gaps in knowledge. In this section, we explore the potential heterogeneity in the benefits of the different problem types across students with different prior performances.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.1">Methods</head><p>To explore whether the effect of problem type varies by prior performance, we added an interaction between students' prior performance and their experimental condition as shown in Equation <ref type="formula">3</ref>.</p><p>The prior performance was calculated as an average of the student's scores from all problems they completed prior to prior to participating in the experiment. In our initial sample, 1,643 students had completed at least ten problems prior to participating in the experiment, and the rest were excluded from the current analysis. This exclusion helps ensure accuracy in inferring the students' mathematical ability, as a limited number of completed problems might not reliably indicate their ability. This exclusion was proportionally distributed across both Fill-In and MCQ conditions-19.96% and 19.72% of students, respectively-ensuring no bias in our estimates due to imbalances in condition-specific sample sizes. In the filtered sample, prior performance scores ranged from 14.29% to 100% with a mean of 70.05% and a deviation of 15.02%. The model standardized the prior performance for better interpretability by z-scoring the prior performance scores. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.2">Results</head><p>Table <ref type="table">4</ref> displays the results for Model 5. The main effect (&#120574; 1 ) is the effect of the Fill-In problem for the students who received the average score because scaled prior performance is centered at the mean. The effect of Fill-In on the likelihood of getting the post-test correct for students with average prior performance effect is nonsignificant (&#120574; 1 = 0.13, SE = 0.09, p = 0.156). The interaction between the prior performance and the Fill-In problem set is significant and positive (&#120574; 3 = 0.20, SE = 0.10, p = 0.042). To ensure that this effect was not a spurious product of our number of prior problems completed cut point of 10, we ran model 3 varying outputs ranging from one prior problem through ten, which did not change the significance or direction of the effects. Therefore, the effect of Fill-In problems compared to MCQs appears to depend on the student's prior math ability-especially for high-performing students. Fill-In problems led to better post-test performance, while for lower-performing students, the effect was smaller and possibly negative. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Observations 3253</head><p>We visualize the interaction between Fill-In problems and prior performance in Figure <ref type="figure">4</ref> by plotting the predicted probability of a correct response on the post-test for each post-test attempt by Prior Performance for both Fill-In and MCQ conditions. These probabilities were predicted using Model 3. Notably, the visualization shows a negative effect of Fill-In problems for students with lower prior scores. To test whether this effect is significant, we ran a post-hoc<ref type="foot">foot_1</ref> model based on Model 5 with prior performance low-end centered so that the main effect will be for the effect of Fill-Ins for students with the lowest prior performance scores. The main effect was not significantly significant &#120574; 10 = -0.61, SE = 0.38, p = 0.114). In summary, we have strong evidence that the effect of Fill-In problems is greater for students with higher prior performance compared with students with lower prior performance. However, there is insufficient evidence that the effect of MCQs is negative for students with lower performance. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7">DISCUSSION AND FUTURE WORKS</head><p>Our analysis found that, on average, problem sets with Fill-In problems were more difficult and led to better learning outcomes than MCQs. However, we observed that the benefits of Fill-In problems have certain contextual constraints. The impact of Fill-In problems on learning was only significant for higher-knowledge students, and there is some evidence that MCQs may be more beneficial for lower-knowledge students. In sum, despite the nuances exposed by our analyses, Fill-In problems have a more positive effect on student learning than MCQs.</p><p>One possible explanation for why Fill-In problems were more effective at improving learning than MCQs is that they increase the likelihood that students will participate in productive struggle. Inducing productive struggle has been shown to increase learning <ref type="bibr">[7]</ref>. Often this is done by creating desirable difficulties-such as varying presentation of content <ref type="bibr">[45]</ref>, interweaving knowledge components instead of presenting them sequentially <ref type="bibr">[41]</ref>, spacing content delivery <ref type="bibr">[11]</ref> and retrieval practice <ref type="bibr">[24]</ref>--during instruction and practice. ORP problems in non-mathematical problem-solving settings are often associated with retrieval practices. Although mathematical Fill-In problems don't solely rely on retrieval, they do require the student to generate their answers independent of any prompts and, therefore, might have a similar benefit to retrieval activities. Taken together, the findings of higher difficulty and greater post-test performance caused by Fill-In problems also align with findings of desirable difficulties, which are often associated with lower students' performance, even as these design choices can positively affect learning as measured by distal outcomes <ref type="bibr">[40,</ref><ref type="bibr">43,</ref><ref type="bibr">46]</ref>. Further evidence that the interplay between problem type, difficulty, and learning may be inducing productive struggle can be found in the other differences in student behavior across conditions. In Table <ref type="table">1</ref>, we reported that the students in the Fill-In problems invested more time presumably thinking before taking their first action than those in the MCQs because they perceived the problems to be more difficult. They were also more likely to utilize hints. These differences may indicate that Fill-In problems are producing better learning behaviors in students, which may be underlying causal mechanisms producing the differences in learning outcomes. Future research should study these potential processes by which problem types impact student learning. One way of doing this would be to use multiple mediation analysis, to evaluate causal paths that lead from problem type to mastery demonstration to discern and assess how problem types influence student behaviors which, ultimately, cause differences in learning outcomes.</p><p>Furthermore, the analysis in Section 6 exploring the heterogeneity of the Fill-In effects indicates that not all students benefit equally from Fill-In problems. We observed that students with higher prior performance benefited more than students with lower prior performance. Although this finding implies that students with lower prior performance might benefit more from MCQs than Fill-In problems; however, we cannot make more substantial claims due to the sparsity of students with low prior performance in the data. Despite this uncertainty, our analysis shows impact differentials for Fill-In based on students' knowledge before beginning the activity. There are some plausible explanations for this phenomenon. High-knowledge students may have the ability to learn the concept addressed in the problem sets but may need the challenge of having to produce the answers themselves without the MCQ options to truly benefit from the activity. Alternatively, lower-knowledge students may benefit from the options in the MCQs but are less likely to learn the concepts well enough to transfer their knowledge to problems where they must provide the answer independently. Regardless of the underlying mechanisms behind the penalization effect, the finding provides evidence that LX designers and instructors may have to consider adapting problem types to students' needs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="8">LIMITATIONS</head><p>There are a few key limitations of our work. First, we conducted experiments on two very specific content areas, where we found evidence that content may influence the effect of the problem type on learning. This research should be replicated using different content areas across different subjects to fully understand the heterogeneity of the impact the problem types can have on student learning. A further limitation of the current study is the lack of student demographic information. The CBLP platform we used in this study does not collect personally identifiable information about the students, per the IRB Protocol; thus, we cannot make any advances in understanding the more fine-grained differences within our sample.</p><p>Our work has an experimental design limitation as we only used Fill-In problems for our post-test. Concerns regarding this design limitation are valid, yet, we argue that knowledge, by nature, should be transferable upon mastery and, as such, would be independent of the instrument used during evaluation. While such assumptions regarding transferability can be problematic, the balanced posttest completion rates across conditions indicate that students from both conditions were comfortable with the design of the post-test. However, we feel that using the Fill-In problem is justifiable as Fill-In problems are an accurate measure of student ability. Further exploration using a combination of both MCQs and Fill-In problems would help establish the optimal approach in the design of assignments as the combination of both activities could enhance learning outcomes or, conversely, the switch between problems in the post-test could cause cognitive load leading to higher dropout rates. Similarly, additional work exploring the benefits and drawbacks of other types of close and open-ended activity design would be beneficial to understand their assessment and learning value.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="9">CONCLUSION</head><p>Overall, findings from the present study present causal evidence that problem types influence how and whether students learn. We observed that, on average, students had better learning outcomes when using mastery-based assignments with Fill-In problems compared to MCQs. We also demonstrated the robustness of our findings by evaluating them across various contextual scenarios, i.e., pre-pandemic, pandemic, and summer sessions. We took a comprehensive approach and evaluated the heterogeneity effects of the two methods, where we observed that high-performing students benefited more from Fill-In problems.</p><p>We hope that the findings of this paper can help inform the design of learning experiences on CBLPs, as we provide evidence that problem types have an impact on learning outcomes. While our findings present the potential benefit of using Fill-In problems in designing learning activities, it is important to highlight that different students, as indicated by their prior performance, may require different types of activity design in order to facilitate more effective learning. We believe that LX designers and instructors will benefit from our findings when designing learning and assessment activities where they are continually required to balance the tradeoffs between the use of open and close-ended activities to facilitate learning while assessing student knowledge.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0"><p>Distractors, also referred to as "Lures" in some academic settings, are incorrect answers in a multiple-choice question designed to mislead students away from the correct answer by providing false information.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1"><p>The full model output is available in the supplemental materials.</p></note>
		</body>
		</text>
</TEI>
