<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Online Search Strategies: Undergraduate Biology Students Take Non-Linear Paths to Find Information Related to Disciplinary Biology Questions</title></titleStmt>
			<publicationStmt>
				<publisher>Journal of Science Education and Technology</publisher>
				<date>04/02/2025</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10616160</idno>
					<idno type="doi">10.1007/s10956-025-10220-5</idno>
					<title level='j'>Journal of Science Education and Technology</title>
<idno>1059-0145</idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Dana L Kirkwood-Watts</author><author>Allison M Johnson</author><author>Sarah K Spier</author><author>Gabrielle B Johnson</author><author>Lorey A Wheeler</author><author>Kathleen R Brazeal</author><author>Brian A Couch</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[With the abundance of information available online, it is important to understand how students search for and use the internet for their courses. Though much is known about general internet search strategies, less is known about how students use the internet for disciplinary coursework. This study aimed to characterize the search strategies students use to find information on the internet to help answer open-ended homework questions. Using observations along with a think-aloud protocol, 25 undergraduate biology students were tasked with answering biology questions of varying complexity to determine how they used the internet and if their use depended on question type. Results indicated that students searched for information in highly recursive ways, shifting in a non-linear fashion between thinking about the problem, performing a search, evaluating search results, and developing an answer. We also found that students tended to rely heavily on curated search engine suggestions and results and that students exhibited only subtle differences in how they sought information for factual versus abstract question types. These findings help expand existing frameworks to include a broader range of behaviors that students employ when using search engines to answer discipline-based questions. This study also provides information that instructors can leverage to help students use the internet as a resource to support their learning, such as by developing structured reflections or by modeling productive information-seeking behaviors.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Higher education courses often use assignments completed outside of class time to help students develop their understanding of course content. While students have acknowledged that they use the internet and social media to help with their courses <ref type="bibr">(Affum, 2022)</ref>, little is known about how students use the internet to answer the questions found on assignments. The internet provides a quick way to access a large amount of information, from simple definitions that allow students to better understand concepts to videos, images, and answer forums directly related to the questions being asked <ref type="bibr">(Brazeal et al., 2021;</ref><ref type="bibr">Ford et al., 2003)</ref>. There are several reasons why we need to develop deeper insights into students' assignment-related internet use. Many current college students are digital natives that grew up with the internet. Browsing the internet for information outside of academic purposes has allowed these students to create their own search strategies before entering college <ref type="bibr">(Bartlett et al., 2020;</ref><ref type="bibr">Thompson, 2013)</ref>, which they utilize when doing any type of search <ref type="bibr">(Chang &amp; Im, 2014;</ref><ref type="bibr">Vitvitskaya et al., 2022)</ref>. Though they are well versed in the use of devices and the internet, their comfort does not necessarily translate to using these tools for academic purposes <ref type="bibr">(Hoeber et al., 2019;</ref><ref type="bibr">Mokhtari, 2014;</ref><ref type="bibr">Mussell &amp; Croft, 2012)</ref>. Furthermore, internet use has increased over the last 20 years with improved search engine algorithms, greater accessibility, and increased flexibility, so students are developing and interacting with internet searches in ways that were not previously possible. For instructors to help students use the internet in ways that support learning, we need to examine students' internet-related behaviors in the context of an academic environment, such as a biology course.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Background Theory and Frameworks</head><p>Broadly, strategies used to search the internet can be classified as information-seeking behaviors, which entail behaviors exhibited in the pursuit, evaluation, and use of information for a purpose. Information seeking begins when a person recognizes a gap in their knowledge and ends when they have successfully filled that gap, completed a task, or given up <ref type="bibr">(Hill, 1999;</ref><ref type="bibr">Humbhi, 2022;</ref><ref type="bibr">Kuhlthau, 1991)</ref>. There have been many contributors to theory development for general information-seeking behaviors, including a description of a four-step process of recognizing a need, searching for the solution, finding the solution, and using the information <ref type="bibr">(Krikelas, 1983)</ref>, and another that described a six-stage process <ref type="bibr">(Ellis, 1993)</ref>. Additional work studying faculty researchers added internet-related behaviors (accessing, networking, verifying, and managing) and then grouped all the behaviors into a more concise model with four stagessearching, accessing, processing, and ending <ref type="bibr">(Meho &amp; Tibbo, 2003)</ref>. In this framework, searching is the starting point and includes the act of recognizing a need and starting to search for information. The next phase is accessing, when the person determines if they can access the information they found in order to use it for their purposes. It includes being able to contact a person knowledgeable about the topic or being able to obtain the information from the internet. The third phase is processing where the person evaluates the course material and synthesizes the information to produce their answer, thus allowing them to reach the end stage of the process.</p><p>Information-seeking behaviors are influenced by many factors <ref type="bibr">(du Toit et al., 2022;</ref><ref type="bibr">Mussell &amp; Croft, 2012)</ref>. First, search behaviors can vary based on the type of information the student is seeking <ref type="bibr">(Kinley et al., 2014;</ref><ref type="bibr">Olsen &amp; Diekema, 2012;</ref><ref type="bibr">Weber et al., 2019)</ref>. Second, the expertise and comfort level of the searcher affects their information-seeking behaviors. For example, graduate students seek information differently than faculty based on differences in their underlying knowledge, confidence, and experience <ref type="bibr">(Gordon et al., 2022)</ref>. Third, task complexity influences search behaviors <ref type="bibr">(Kinley et al., 2014)</ref>. For factual tasks, the answer to a question is relatively easy to find with a few searches. In contrast, in response to more complex or abstract tasks that may not have a direct answer, people usually exhibit more comprehensive search behaviors, including conducting a greater number of searches, incorporating more words per search, and using more complex search terms <ref type="bibr">(Ghosh et al., 2018;</ref><ref type="bibr">Kinley et al., 2014)</ref>.</p><p>Building on the existing models, a recent study characterized the specific behaviors that students exhibit as they search for general information online <ref type="bibr">(Hinostroza et al., 2018)</ref>. This research identified a variety of actions that students take and aligned these behaviors to four main processes: defining and understanding the problem; using the search engine; scanning, evaluating, and selecting webpages; and processing and integrating into an answer. While students may engage in consistent search patterns, these researchers found that students also adjust their behaviors based on the task requirements and their emerging information needs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Rationale for Current Study</head><p>While students can develop their search strategy skills as they advance in their digital literacy <ref type="bibr">(Atoy et al., 2020)</ref>, students often have poor search strategies and as a result do not learn effectively from their searches <ref type="bibr">(De Simone et al., 2021)</ref>. However, existing research has focused on non-academic search tasks, which may elicit different search strategies than academic tasks. For non-academic tasks, the focus is on quick information rather than information accuracy, and studies have characterized topics such as whether the information is accessed via search engine or social media and which sources are used <ref type="bibr">(Bartlett et al., 2020;</ref><ref type="bibr">du Toit et al., 2022)</ref>. For academic tasks, the focus shifts to gathering correct and reliable information; however, students tend to have trouble switching strategies between academic and non-academic searches <ref type="bibr">(du Toit et al., 2022)</ref>.</p><p>There have been some academic studies that cover subject specific areas, such as English composition <ref type="bibr">(Olsen &amp; Diekema, 2012)</ref>, nutrition <ref type="bibr">(Adamski et al., 2020)</ref>, and preschool education courses <ref type="bibr">(Kao &amp; Chien, 2017)</ref>, but very few studies relate search behaviors to learning outcomes (e.g., <ref type="bibr">Nagel et al., 2020;</ref><ref type="bibr">Safdar et al., 2020;</ref><ref type="bibr">Weber et al., 2019)</ref>. A few studies examine students' information-seeking behaviors, finding that study approaches, such as being actively interested in the topic, affect search strategies by promoting advanced search techniques (e.g., Boolean searches; <ref type="bibr">Ford et al., 2003)</ref> and describing how students click through webpages and modify their searches to find relevant information <ref type="bibr">(Skripchuk et al., 2023)</ref>. Others show that higher performing students tend to do more advanced searching and that task complexity affects search behaviors <ref type="bibr">(Ghosh et al., 2018;</ref><ref type="bibr">Weber et al., 2019)</ref>. While these studies provide a foundation for understanding, they compare search strategies to study strategies or course performance, rather than directly relating search strategies to answer generation or correctness. We need more direct information on how undergraduate students use the internet to complete the types of questions commonly found in course assignments.</p><p>Understanding how students access the internet and apply information can help instructors guide their students to use this resource productively in their courses. Our study sought to understand how undergraduate biology students use the internet for academic purposes by investigating how students find information online to answer subject specific (i.e., biology) questions. We looked specifically at how they design their searches and how their information-seeking behaviors relate to their answers. We also examined different task types to determine how question complexity affects search strategies. The following research questions guided the overall study:</p><p>RQ1: What behaviors do undergraduate students from introductory level biology courses exhibit while searching for the answer to biology questions using the internet, and do these behaviors vary based on task complexity? RQ2: How do students from introductory level biology courses design their searches to find the answer to biology questions, and is there variation based on task complexity? RQ3: To what extent do aspects of a student's search relate to the correctness of their answer to biology questions?</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Methodological Approach and Interview Procedure</head><p>We conducted a descriptive study to determine how undergraduate students from an introductory biology course use the internet to answer questions they might see in a homework assignment (see Supplemental Materials Figure <ref type="figure">S1</ref> for an overview). We used various forms of triangulation throughout the data collection and analysis process to support the credibility of our qualitative findings <ref type="bibr">(Twining et al., 2017)</ref>. In particular, we used both a theoretical framework <ref type="bibr">(Meho &amp; Tibbo, 2003)</ref> and a conceptual framework <ref type="bibr">(Hinostroza et al., 2018)</ref> to inform our interview protocol and subsequent coding. We synthesized information from multiple data sources, and multiple researchers classified, evaluated, and analyzed student search behaviors and biology answers using qualitative and quantitative approaches as a means to summarize student approaches and address associated research questions.</p><p>The biology questions used in this study were designed to target concepts just beyond the typical conceptual outcomes for an introductory biology course to help ensure that the questions required the students to use the internet. Similar to previous information-seeking studies, we developed eight scenarios that each contained two distinct types of tasks: factual and abstract <ref type="bibr">(Kinley et al., 2014)</ref>. We used questions from a textbook question bank as inspiration for the topics, rewriting them with a scenario to provide a context for the subsequent tasks. For each scenario, we developed tasks classified as factual and abstract. The factual task asked for definitions or short descriptions of processes and could be readily found through web searches. The abstract task required the student to formulate an answer using multiple sources of information. The abstract tasks were designed to challenge students to integrate and apply, rather than just find information. An example scenario stated, "You are studying a protein that promotes cell division in human skin cells. You discover that this protein often becomes attached to another protein called ubiquitin." The factual question followed: "What is the normal function of ubiquitin in cells?" Once the student answered the factual question, the abstract question appeared: "What do you predict would occur in a cell line that has been modified in such a way that the protein you are studying cannot be attached to ubiquitin?" A full list of scenario questions and tasks can be found in Supplemental Materials Table <ref type="table">S2</ref>.</p><p>Interviews were conducted during the spring 2020 and fall 2020 semesters via videoconference (Zoom) to allow for recorded observation of the student's screen, along with recorded communication between researcher and participant. The participants worked through a Qualtrics survey containing the biology scenarios and questions along with text boxes to enter their answers. During the interview, to maintain confidentiality for the interview recording, the participant was given a participant code and instructed to change their Zoom name to this code. Once this was changed, the participants shared their screens and recording of the session began. Then, the researcher went over the consent form with participants and introduced the interview process. This included instruction to proceed with answering the questions as they would if it were a homework activity for a biology course, such as using their preferred browser, their typical search strategies they employ when they search for biology information, and their typical behaviors when answering questions. The study also utilized the think-aloud method for students to verbalize their mental processes as they completed the tasks <ref type="bibr">(Ericsson &amp; Simon, 1980)</ref>. In the event the student stopped describing their methods, the researcher asked probing questions and reminded the student to talk through their process.</p><p>With respect to method triangulation, our data were in the form of interview transcripts, observations from search behaviors, and written question responses. Zoom recordings were transcribed using the Microsoft Word Transcribe function in Office365. The researchers watched the interviews and annotated the transcripts with behaviors that occurred during the recording that were not captured in the transcripts. Items added to the transcript included the search string the student used, the order in which the search strings were used, the sequence of events the student undertook to obtain their answer, the websites used in formulating the answer, and behaviors from the interview that were otherwise not captured on the transcript.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Research Context and Participant Sample</head><p>Our target population for the study was students in introductory molecular and cellular biology courses for science majors. Students were invited to participate through an online announcement in their associated biology courses and were compensated $40 for their time. A total of 25 students were recruited from three institutions: 17 from a public 4-year research university and eight students from two different associate's institutions. Student demographic characteristics are noted in Table <ref type="table">1</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Development of the Codebook and Coding of the Interviews</head><p>Using a starting codebook derived from <ref type="bibr">Hinostroza et al. (2018)</ref>, we conducted a priori coding on one question of one interview representative of responses we saw during the data collection process. The a priori coding allowed us to use the existing codebook to identify and categorize items in our data. Codes were applied to words, phrases, and sentences within the transcripts. After this initial coding, investigator triangulation occurred among five authors who discussed and agreed upon adjustments to the codebook. These changes were applied, and a second round of coding was done by the same five authors, this time with a set of two questions from two different interviews. We continued this process of coding, discussing, and adjusting the codebook for a total of seven rounds of two interview questions each, or approximately 20% of the dataset, until we came to agreement. One author then applied the codes to the remainder of the dataset, consulting the other four authors when ambiguities in transcripts arose. The final coding of the transcripts was completed in Dedoose, then downloaded, and analyzed in R and Microsoft Excel (Dedoose Version 9.0.17, 2021; RStudio <ref type="bibr">Team, 2020)</ref>.</p><p>For certain analyses, we collapsed the codes into condensed code groups with conceptual and behavioral similarities with specific attention to distinguishing codes based on the originality of student thought patterns. Codes were collapsed into nine groups, which captured the broad behaviors students used during the interviews. Additionally, groups were assigned to behaviors that were of specific interest to the study, such as copying and pasting either the question or the answer, visiting websites, or using only the search results page.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Data Analysis</head><p>Code frequency was used to see how often codes appeared within the student interviews. This revealed the prevailing actions performed by students throughout the interview. The frequency of codes was calculated by totaling the number of each code a student had during their interview. Since not all students were able to answer the complete list of questions during the interview, the number of each code that occurred during the interview was divided by the number of tasks the student answered. These individual student values were averaged across students to get the overall code frequency.</p><p>Sequence plots and Markov models used the condensed groups and showed the sequence of events that occurred while the student was answering the biology questions and the probability of moving from event to event. The sequence plots allowed us to see how each question was being answered and to compare search sequences for factual and abstract tasks. The Markov models allowed us to compare the probabilities of the students moving across events (groups) between the two task types. The sequence plots were created using the TraMineR package <ref type="bibr">(Eva &amp; Futing, 2014;</ref><ref type="bibr">Gabadinho et al., 2011)</ref>. The Markov models were produced using markovchain <ref type="bibr">(Spedicato, 2017)</ref>. We analyzed a variety of search query characteristics to determine if there were differences across task types. We calculated the number of searches per task for each student and then averaged these values across students. Word count was calculated to address task complexity, with an expectation that higher task complexity would elicit more words per query <ref type="bibr">(Ghosh et al., 2018)</ref>. These values were calculated by taking the average number of words per query for each student, and then averaging across students. Edit distances have also been used as a measure of task complexity <ref type="bibr">(Ghosh et al., 2018)</ref>. The edit distance is the difference between two consecutive strings determined by the number of edits required to get from string 1 to string 2. It quantifies string similarity by calculating the minimum number of editing operations (insertions, deletions, substitutions) required to change string 1 into string 2. The more editing operations required indicate that the two strings are more dissimilar. The average edit distance was calculated by taking the average edit distance per task for each student and then averaging across students. Finally, we categorized searches based on if the query was a direct copy/paste from the original: if the query was a paraphrase in the form of a keyword search, question, or sentence or if the student did not perform a search at all. The number of searches, word count, and types of search queries were categorized and counted in Excel and analyzed using the R package stringR <ref type="bibr">(Wickham, 2002)</ref>. Edit distances and the percentage of the longer string were calculated in Excel using a macro obtained from Exceldemy.com.</p><p>We investigated steps that students took after each search query, again comparing across task types. First, we quantified the proportion of times following a search that a student stayed on the results page versus clicked on a website. We then categorized websites into groups based on similarities in content and publisher, such as scientific, medical, learning, and university sources. Scientific websites included Nature, NCBI, ScienceDirect, Springer, and any other website that primarily publishes peer-reviewed journal articles. Medical websites included websites with medical information, including hospital websites and websites specifically designed for disease education. Medical websites were distinguished from health websites, which included health information such as would be found on a website for medical conditions or a professional website from a doctor. Learning websites included websites typically used for student learning, such as Khan Academy and Lumen Learning, and were distinguished from university websites, such as any website ending in.edu.</p><p>Student question responses were scored incorrect/correct on a 0/1 scale. Dual coding of every response was achieved by having five authors each score one or two questions and one author score all the questions. Any scoring differences were discussed to consensus. A generalized linear mixed model was performed with answer correctness as the binomial outcome variable, the items of interest as fixed effects (search engine, number of searches, search type, website use, and task type), and student and question as random effects. This was produced using the R package lme4, with the link logit and family as binomial <ref type="bibr">(Bates et al., 2015)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>RQ1: What Behaviors do Undergraduate Students from Introductory Level Biology Courses Exhibit While Searching for the Answer to Biology Questions Using the Internet, and do These Behaviors Vary Based on Task Complexity?</head><p>We started with a codebook from a previous study and adjusted to accommodate our observed student behaviors <ref type="bibr">(Hinostroza et al., 2018)</ref>. Table <ref type="table">2</ref> shows an abbreviated codebook; a full descriptive codebook can be found in Supplemental Materials Table <ref type="table">S3</ref>. Noted differences from the original codebook included expanding the categories to include instances where students used background knowledge to help in their search, adding codes to capture newly observed student search behaviors, and introducing codes that reflect changes in internet search functions such as "Google Suggests" and "People Also Ask." We also found that students tended to return to websites and previous searches throughout the interview, and therefore, we added codes to the Evaluates Results category to reflect these behaviors. Finally, during the interviews, we noticed some students expressed uncertainty either about the question or about their answers, leading us to add codes to reflect when they qualified their answers. Table <ref type="table">3</ref> provides a representative example of a coded interview transcript.</p><p>After applying the codebook, we examined the frequency of codes used within our study (Fig. <ref type="figure">1</ref>), or how often a given student used the code for a particular task. Notably, when students designed and conducted their searches, they often used Google's People Also Ask function. Consistent with previous studies, students did not typically use advanced search techniques, such as Boolean searches, navigating directly to a particular website, or using quotes to refine their search results <ref type="bibr">(Lowe et al., 2020)</ref>. Though they used websites to help answer questions, they did not often explicitly consider the usefulness or validity of the source they used.</p><p>When the students clicked on websites, it was a website found in the first third of the results, consistent with previous studies showing that first results are most often used <ref type="bibr">(Nagel et al., 2020)</ref>. Finally, students tended to paraphrase answers rather than copying answers or using their own ideas.</p><p>Sequence plots showed how individuals moved between different behaviors (condensed groups) as they developed their answers (Fig. <ref type="figure">2</ref>). The graphs on the left display the raw patterns that occurred when students answered the tasks, what actions the students performed, and in which order. This allowed us to visualize the substantial variance that occurred as students addressed the various questions, while also capturing the iterative and sequential nature of events. In the right-hand plots, the data are disconnected from each student/question and instead show the proportion of activities occurring for each step in a generalized sequence. The first step that most students took was to read the question. However, a small portion of students appeared to be thinking first without apparent reference to the question, while another portion engaged in rote behaviors, such as directly performing a search using a recurrent strategy or copying the question. These graphs show subtle differences in the sequences based on task type, such as in the second step where students were more likely for a factual task to perform a search (0.56, compared to 0.41 for abstract) or copy the question (0.18, compared to 0.13 for abstract) and were more likely for an abstract task to think about the problem (0.24, compared to 0.14 for factual) or write an answer (0.21, compared to 0.08 for factual).</p><p>Markov models showed the probability of students transitioning between different states or condensed code groups (Fig. <ref type="figure">3</ref>). These patterns differed in some ways between the two task types. To examine these differences, we took the difference between the transition matrices for the Markov models of factual and abstract questions (Table <ref type="table">4</ref>). For factual questions, the most notable transitions were that students more often read the question or thought about the problem immediately before performing a search, copied the question and then evaluated the results, and went from using an image to writing an answer. For abstract questions, students more often went from reading the question or thinking about the question directly to writing an answer and from copying the question to using a webpage. With respect to recursive steps, students addressing factual questions more often returned to perform another search or evaluate results. For abstract tasks, students more often went from using an image to re-reading the question, performing another search, or evaluating results. Students also had greater tendency to go from copying an answer to using a webpage for abstract tasks, although it should be noted that copying answers was a relatively uncommon code.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>RQ2: How do Students from Introductory Level Biology Courses Design Their Searches to Find the Answer to These Biology Questions, and is There Variation Based on Task Complexity?</head><p>We next explored how students designed their search queries (i.e., the specific entries students submit to the search engine). The most common search engine was Google (80% of students), with other search engines represented including Bing, Ecosia, and Ocean Hero. This is consistent with global search engine usage. Students performed an average Notes a guess at the answer Expresses uncertainty about answer of 1.5 queries per task, with no difference between the factual and abstract task types (Fig. <ref type="figure">4a</ref>; t = -0.28, p = 0.78).</p><p>In evaluating task complexity, the overall word count per query was 6.2, but this differed between factual (5.7 words) and abstract (7.0 words) tasks (t = -3.48, p &lt; 0.001). Edit distances indicated that when a student altered their query, they adjusted more than half the query, regardless of task (Fig. <ref type="figure">4b</ref>; t = 1.18, p = 0.24).</p><p>In addition, we examined how students designed their queries and if the design differed based on task type (Fig. <ref type="figure">4c</ref>). We found that there were distinct ways students designed their searches: copying and pasting from the original question, using keywords, writing their own question related to the original, writing a sentence related to the original, or not conducting a search. Overall, both factual and abstract tasks elicited the same types of searches with no statistical significance between them. The most common search type for both tasks was a question format (i.e., asked the search engine a question other than the original question). While it did not reach significance (t = -1.94, p = 0.07), there was a slight tendency for students to start the answer to an abstract task without a search. Some of these students recognized that the answer to the abstract task may not be easily found on the internet and opted to try to answer without a search. Others indicated that they found enough information from a previous search to attempt the answer.</p><p>Finally, we looked at where students found their information to answer the task, on the results page or on a website, and if this differed based on task complexity (Fig. <ref type="figure">5a</ref>). Students tended to stay on the results page more often when they were answering a factual task than an abstract task (t = 2.34, p = 0.02). Conversely, students were more likely to use a website when answering an abstract task (t = -2.28, p = 0.03). When they did click a website, however, they predominantly chose a scientific website, regardless of question type (Fig. <ref type="figure">5b</ref>). In cases where students selected a non-scientific website, some noted that scientific websites have more complex answers than they can understand and others recognized that there is a difference in the types of websites but did not spend time dwelling on it. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Considers source</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Reads webpage</head><p>Rereads question Writes answer</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Uses webpage Rereads question Paraphrases answer</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>RQ3: To What Extent do Aspects of a Student's Search Relate to the Correctness of Their Answer to Biology Questions?</head><p>We wanted to determine if search engine, number of searches, search type, website use, or task type related to how students answered the questions. We estimated a generalized linear mixed effects model with student and question as random effects and included these search variables as predictors to see if any of them were associated with a correct answer (Table <ref type="table">5</ref>). The only significant predictors for a correct answer were search engine and task type. Students were more likely to get the answer correct if they were using Google as their search engine or answering a factual question.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head><p>The purpose of this study was to characterize internet search strategies used by students to find biology information related to a course context. We found that, though students had some progressive strategies, they tended to have a more iterative process. Relating our results to Hinostroza et al.'s (2018) framework, we found that biology information seeking aligned with the four areas of that framework: defining and understanding the problem, using the search engine; scanning, evaluating, and selecting webpages; and processing and integrating into an answer. Here, we provide a summary of the salient behaviors that students exhibited in each phase. This graph is arranged first by codebook section, then within codebook section by increasing code use. Error bars represent standard error</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Defining and Understanding the Problem</head><p>In the initial phase, students read the question and tried to understand the information-seeking task. They also thought about the problem and talked through some avenues to answer. Though they tended to think about abstract tasks more often than factual tasks, there was only slight variation in this behavior. This starting phase was a relatively short phase in which the students decided if they needed to conduct a search to answer the question. Those students that Fig. 3 Markov models for A factual and B abstract questions. The nodes represent condensed code groups, with the numbers in the nodes reflecting the likelihood of this being the starting step. The numbers on the gray arrows reflect the probability of students transitioning to a different state from the state they were in previously. Gray arrows on the top half reflect transitions from a left node to a right node and on the bottom half reflect transitions from a right node back to a left node. Transitions with probabilities less than 0.1 are suppressed for clarity</p><p>answered the question without a search were not undertaking any behaviors related to information seeking. A common misconception is that if the internet is available, students will use it even if they know the information <ref type="bibr">(Ferguson et al., 2015)</ref>; however, students did not always reference the internet. To elicit more information-seeking behaviors, the task should be outside the scope of a student's current or perceived knowledge.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Using the Search Engine</head><p>When in the searching phase, students copied the question, paraphrased the question, or conducted a search tangentially related to the original question. They tended to only do one or two searches per task. When they did more than one search, more than half of the search query was rephrased. These findings align with the hypothesis that students tended to do what they were accustomed to doing and stayed within their habit of conducting only one search and relying on the first results, even within the realm of academic information seeking <ref type="bibr">(Olsen &amp; Diekema, 2012)</ref>. Students also relied heavily on the search engine's proposed queries, which suggests that students may not have been thinking as critically about the question they were asked. Behavior patterns did not vary much based on whether the student was answering a factual or abstract question. Though previous literature states that information-seeking behaviors are dependent on the complexity of the task being performed, in seeking biology information, we did not see major differences in the searching phase of the information-seeking task <ref type="bibr">(du Toit et al., 2022)</ref>. This may be because the tasks were not different enough from each other or it could be that students did not modify their search strategy based solely on task type and may have had other reasons to adjust their behaviors, such as how an assignment will be used or graded in the course.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Scanning, Evaluating, and Selecting Webpages</head><p>During this phase, students evaluated the results and decided if they would click a website or image. We saw variation in strategy based on task type during this phase. Students tended to stay on the results page more for factual questions, whereas they tended to click a website more for abstract questions. This behavior reflected search engine algorithms that are designed for quick, definition-type answers, thereby allowing students to remain on the results page to answer a factual task. Conversely, these differences point to students' need for a deeper examination of the concept to answer the abstract task. Regardless of task, students typically selected scientific-based websites related to the subject area when they chose to click on a webpage. Students did not seem to extensively question the validity or accuracy of the information found in search results or scientific websites, which may stem from search engines prioritizing heavily accessed sources and from students entrusting the search engine with finding mainstream and acceptable information.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Processing and Integrating into an Answer</head><p>During this phase, students started producing their answers and made decisions about whether they needed to return to any of the previous phases to continue information seeking.</p><p>The Markov models indicated how students often returned to perform a search, evaluate results, use a webpage, or use an image. Between task types, factual tasks tended to instigate recursive moves as students were generating their answers (i.e., from copying or writing an answer back to performing a search or evaluating results). This result suggests that factual tasks may lead students to check or elaborate their answers by re-engaging with the information-seeking process. Conversely, abstract tasks had recursive moves prompted by using an image (i.e., from using an image back to reading the question, performing a search, or evaluating results), which implies that students in these cases may be seeking additional information based on what they found within an image. Thus, for biology questions, factual questions may lead students to search modes that prioritize reiterating information while abstract questions encourage students to find and then clarify information. We wanted to see if any search characteristics related to how students answered the questions. The only predictors of answer correctness were search engine and task type. Students were more likely to get the answer correct if they used Google as their search engine and more likely to get factual questions correct. This result likely stems from the marked ability of search engines to display factual results. When looking for a definition of a word or description of a process, search engines will display as close to a definition or description as possible, many times from a dictionary or Wikipedia. Conversely, if the query is more abstract, search engines may have more trouble displaying a direct answer. Though this study showed Google as being predictive of answer correctness, we observed that other search engines presented students with similar results as Google, suggesting some additional difference in how students translate information into answers.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Conclusions</head><p>We explored the information-seeking behaviors undergraduate biology students employ as they search for the type of biology information they would use for a course assignment. With respect to RQ1, we found that students follow an iterative process while searching for information about biology topics <ref type="bibr">(Gordon et al., 2022)</ref>. Our Markov models extend the current understanding of student search behaviors by providing a visual representation of the complexity behind search behaviors, with students shifting across the code groups in a highly recursive manner suggestive of a search space rather than a specific search path. We also found that students rely heavily on the search engine, such as by using the Google Suggest dropdown search and the People Also Ask questions, by generally selecting the first result, and by infrequently considering the accuracy of the source <ref type="bibr">(Skripchuk et al., 2023)</ref>. This suggests that students bestow a certain degree of trust to search engines, potentially accumulated over years of conducting searches that produce satisfactory results. Our study also found mostly similar behaviors across task types, with some differences in how frequently students engaged in particular forward and reverse pathways. For RQ2, we found that students averaged one and a half queries per question regardless of task complexity, again showing a willingness to rely upon the search engine's results. However, for more complex tasks, students used more words per query and tended to click websites more often. These results suggest that the relationship between task complexity and search string length is independent of education contexts <ref type="bibr">(Ghosh et al., 2018)</ref>. For RQ3, previous research did not directly compare search strategies with answer correctness, and we found that the predictors of answer correctness were limited to the search engine used and the type of question being answered.</p><p>As with any observational study, participants may have altered their behaviors from what they would do if they were not being observed. One student indicated that they would approach their tasks differently depending on how the questions were graded within a course, so future research should explore how grading affects search strategies. Some codes were inferred from the interviews and therefore could represent undercounts, such as when a student read the question. At times, it was unclear if the student was reading the question to start off; however, students did tend to go back and reread the question, which was more evident from their behaviors.</p><p>Possible explanations of why we saw only subtle variation in information-seeking behaviors between task types may be in how we designed the questions. Though they aligned with the "factual" and "abstract" criteria and the question scenarios were given in a random order, the pairs (factual and abstract) concerning the same content appeared together and the student always saw the factual question first and the abstract question second. This may have also contributed to our observation that the students tended to answer the abstract question more often without a search. Follow-up studies could separate the factual and abstract questions to determine if there are more differences in search strategies when the student does not first receive a factual question.</p><p>The generalizability of these results may be limited to introductory undergraduate students seeking science content for their courses. Because the way people seek information varies not only with the task being done but also their prior knowledge, it is reasonable to suggest that an upper-level undergraduate student or graduate student would seek information differently than an introductory student <ref type="bibr">(Adamski et al., 2020;</ref><ref type="bibr">Gordon et al., 2022)</ref>. While our sample included first-year and non-first-year students, these students had similar biology coursework and so they are not ideal for studying how student behaviors may change throughout college. Instead, we propose that comparing 100-level versus 400-level students represents a promising avenue for understanding shifts that may occur in information seeking during college. This work highlights opportunities for instructors to help increase students' out-of-class engagement with their homework. Information-seeking behaviors are an important aspect of higher education. Though students come to college already knowing how to search the internet, these behaviors may not be sufficient for the academic tasks to which they will be exposed. Information seeking can be classified as a science process skill to be nurtured and honed. Furthermore, this skill potentially transfers across disciplines and helps students become proficient information consumers.</p><p>Our study showed that students engage in four overarching behaviors as they search for biology information. We found that information seeking is not a linear process but rather takes complex and iterative paths. This study provides information that instructors can leverage to help students use the internet as a resource to support their learning. For example, a structured reflection could be developed to help students improve in all phases of information seeking. This reflection could present our abbreviated codebook (Table <ref type="table">2</ref>) as a way for students to see the many possible actions that occur during information seeking. The Markov diagrams (Fig. <ref type="figure">3</ref>) could provide a way for students to see the recursiveness of search patterns. In both cases, students could reflect on their current approach, evaluate the strengths and limitations of their strategy, and identify improvements that align with deeper learning. Furthermore, instructors can help students practice with tasks of varied complexity and implement a broader range of information-seeking behaviors by modeling productive behaviors in response to factual and abstract tasks. Instructors could provide specific guidance in each search phase (i.e., thinking about the problem, performing searches, evaluating results, and developing an answer) and give insights into how students might choose specific actions and sequences within their search. In these ways, instructors can help students develop their process skills and become more deliberate about how they use the internet as a support for their learning.</p></div></body>
		</text>
</TEI>
