<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Don’t Mention It? Analyzing User-Generated Content Signals for Early Adverse Event Warnings</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>09/01/2019</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10179267</idno>
					<idno type="doi">10.1287/isre.2019.0847</idno>
					<title level='j'>Information Systems Research</title>
<idno>1047-7047</idno>
<biblScope unit="volume">30</biblScope>
<biblScope unit="issue">3</biblScope>					

					<author>Ahmed Abbasi</author><author>Jingjing Li</author><author>Donald Adjeroh</author><author>Marie Abate</author><author>Wanhong Zheng</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Please scroll down for article-it is on subsequent pagesWith 12,500 members from nearly 90 countries, INFORMS is the largest international association of operations research (O.R.) and analytics professionals and students. INFORMS provides unique networking and learning opportunities for individual professionals, and organizations of all types and sizes, to better understand and use O.R. and analytics tools and methods to transform strategic visions and achieve better outcomes. For more information on INFORMS, its publications, membership, or meetings visit http://www.informs.org]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Product-related adverse events can have profound monetary and societal implications in various industry contexts. For instance, adverse pharmaceutical drug reactions are responsible for between 3% and 12% of all hospital admissions <ref type="bibr">(Ritter 2008)</ref>, resulting in millions of hospitalizations and more than 100,000 deaths annually <ref type="bibr">(Sarker et al. 2015)</ref>. The pharmaceutical drug Pradaxa alone has caused 9,000 hospitalizations, 1,000 deaths, and $650 million in lawsuit settlements since 2014 <ref type="bibr">(Thomas 2014</ref><ref type="bibr">, Colella-Walsh 2019)</ref>. Similarly, in the automotive industry, Toyota recently settled lawsuits totaling $3.4 billion and $2.3 billion for inadequate rust protection on their trucks, and the unintended acceleration "sticky pedal" fiasco, respectively <ref type="bibr">(Schweinsberg 2012</ref><ref type="bibr">, Fredericks 2014)</ref>. Traditional adverse event reporting mechanisms have involved formal reporting systems that feed into online databases. Examples include the Adverse Event Reporting System (FAERS) of the U.S. Food and Drug Administration (FDA) and the National Highway Traffic Safety Administration's (NHTSA) safety issues database. Such databases constitute an invaluable data source for postmarket surveillance. However, many studies have noted the limitations of overreliance on a single channel-most notably, limited coverage of the broad set of adverse events encountered by a diverse consumer population <ref type="bibr">(Forster et al. 2012)</ref>.</p><p>With the rise of big data analytics <ref type="bibr">(Agarwal and Dhar 2014)</ref> and greater impetus on broader postmarket surveillance, the Voice of the Customer (VoC) has emerged as an important source of information for understanding consumer experiences and identifying potential issues <ref type="bibr">(Zabin et al. 2011</ref><ref type="bibr">, Boynton 2013)</ref>. This is partly due to increased quality, volume, and timeliness of available VoC information, which encompasses various user-generated content channels including social media, search queries, consumer reports, etc. One of the biggest VoC use cases remains managing risk <ref type="bibr">(Zabin et al. 2011</ref><ref type="bibr">, Browne 2015)</ref>. For instance, in 2014, for the first time ever, the FDA received more adverse drug reports from consumers than from healthcare professionals-and the volume of customer search queries was orders of magnitude higher <ref type="bibr">(White et al. 2013)</ref>. Similarly, social media channels offer great potential for adverse event detection for a myriad of products <ref type="bibr">(Abrahams et al. 2015</ref><ref type="bibr">, Sarker et al. 2015)</ref>. Relevant stakeholders interested in leveraging VoC include regulatory agencies, product manufacturers, consumer advocacy groups, and financial investment firms. In such organizations, risk management groups are increasingly interested in working with their information technology (IT) teams to develop robust VoC listening platforms <ref type="bibr">(Fenwick et al. 2011</ref>) capable of identifying adverse events faster and more accurately, resulting in favorable economic and humanistic outcomes.</p><p>However, two key challenges have impeded the success of VoC listening platforms. First, the existing body of knowledge has leveraged diverse sets of channels, adverse event types, and modeling methods, resulting in varying results and diverging conclusions regarding the viability and efficacy of various online user-generated channels and accompanying modeling methods <ref type="bibr">(Schmidt-Subramanian et al. 2014</ref><ref type="bibr">, Sarker et al. 2015)</ref>. As <ref type="bibr">Davies (2016, p. 1)</ref> noted, "A myriad tools and techniques can be applied to a VoC program. This complicates the tasks of investment prioritization and feedback alignment." Second, many existing detection methods rely on "mention models" that have low detection rates, have high false positives, and fail to detect adverse events in a timely manner, rendering them less useful in real-world risk management contexts <ref type="bibr">(Adjeroh et al. 2014</ref>). Consequently, "a key stumbling block for many VoC initiatives" is the lack of meaningful, actionable insights <ref type="bibr">(Davies 2016, p. 8)</ref>. Presently, risk management and monitoring groups, and IT teams that support such groups, are lacking guidelines regarding many key questions such as the following <ref type="bibr">(Schmidt-Subramanian et al. 2014</ref><ref type="bibr">, Davies 2016)</ref>: "Which channels should we be integrating into our listening platform?" "Which types of detection methods are best suited for our event types?" "How can we design listening platforms that are practical and valuable in our monitoring contexts?" There remains a need to examine the efficacy of various VoC channels for IT applications with implications for consumer safety <ref type="bibr">(Agarwal et al. 2010</ref><ref type="bibr">, Abrahams et al. 2015)</ref>. Furthermore, recent studies have underscored the need for more robust detection methods applied to these channels that can serve as decision aids for monitoring teams. The two main research questions we seek to answer are as follows:</p><p>1. How effectively can various VoC channels be used to detect different types of adverse product events using state-of-the-art signal detection methods?</p><p>2. What are the relevant interactions between channels, event types, and modeling methods, and what are their implications for the design of VoC listening platforms?</p><p>To tackle these questions, following the information systems (IS) design science approach, in this research note we propose a framework for examining key design elements for VoC listening platforms. As part of our framework, we also develop a novel heuristic-based method for detecting adverse events. We evaluate our framework and method on two large test beds, each encompassing millions of tweets, forum postings, and search query logs pertaining to hundreds of adverse events related to the pharmaceutical and automotive industries. The results shed light on the interplay between user-generated channels and event types, as well as the potential for more robust event modeling methods that go beyond basic mention models. More specifically, the results from our analysis framework reveal that user-generated content channels can facilitate timelier detection of adverse events: on average, two to three years earlier than commonly used regulatory databases. The inclusion of negative sentiment polarity in the models can further reduce false-positive rates across all three channels. Additionally, we find social media channels provide higher detection rates but lower precision than search-based signals. In the context of more explicit/salient events, search and web forum channels are timelier than Twitter. Furthermore, certain event types such as drug-related product recalls are more challenging to detect using user-generated content channels. Whereas most existing mention models detect less than half of all events earlier, with false-positive rates over 75%, the proposed heuristicbased method attains markedly better results-with earlier detection rates of 50%-80% and far fewer false positives across an array of VoC channels and event types. The heuristic method is also well suited for signal fusion across channels.</p><p>Our note makes several key contributions to research and practice. We contribute to the emerging IS body of research developing novel analytics capabilities with important business and societal implications (e.g., <ref type="bibr">Shmueli and Koppius 2011</ref><ref type="bibr">, Chen et al. 2012</ref><ref type="bibr">, Bardhan et al. 2015</ref><ref type="bibr">, Brynjolfsson et al. 2016)</ref>. From a design science perspective, our contributions include a holistic framework for analyzing key design elements pertaining to VoC listening platforms, as well as a novel heuristic-based event modeling method. Our framework unifies and expounds upon insights and key design elements previously examined in a disparate manner, affording opportunities to better understand the interactions between channels, event types, and modeling methods. The proposed event modeling method offers robust detection capabilities that are largely channel and event agnostic, across multiple industry contexts, thereby shifting the detection paradigm away from the status quo, underperforming mention models.</p><p>Finally, our research has managerial implications for various practitioner groups. The framework and results can offer guidelines for risk management groups and IT teams tasked with developing and operating VoC listening platforms. By incorporating provisions for key monitoring objectives and constraints such as timeliness, detection, and false-positive rates, the framework is well suited for use by several stakeholders, including regulatory agencies, manufacturing firms, and advocacy groups. Furthermore, the proposed method provides markedly better detection capabilities, making VoC listening practical and valuable.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Proposed Framework</head><p>Organizations broadly recognize the importance of listening to VoC, with risk management cited as a primary use case <ref type="bibr">(Zabin et al. 2011</ref><ref type="bibr">, Abrahams et al. 2013)</ref>. However, the percentage adoption of robust VoC listening platforms and firms' perceived capability maturity of their platforms have both been problematic <ref type="bibr">(Browne et al. 2015)</ref>. A core issue is that presently, "knowledge of designing, building, integrating, and modifying" effective VoC listening capabilities remains low <ref type="bibr">(Davies 2016, p. 5)</ref>. There is a need for design frameworks that can bridge the gap between why organizations deploy VoC listening platformsnamely, to integrate appropriate channels and derive valuable insights-and actual outcomes <ref type="bibr">(Zabin et al. 2011)</ref>.</p><p>Design science provides guidelines for the development of IT artifacts, including constructs, models, methods, and instantiations <ref type="bibr">(Hevner et al. 2004)</ref>. Several prior studies have utilized a design science approach to develop business intelligence and analyticsrelated IT artifacts, including frameworks, methods, and instantiations (e.g., <ref type="bibr">Lau et al. 2012</ref><ref type="bibr">, Provost et al. 2015)</ref>. In this note, we employ the design science approach to develop our proposed analysis framework for designing VoC listening platforms.</p><p>When creating IT artifacts in the absence of sufficient guidelines, the design science literature suggests that kernel theories may help govern the development process <ref type="bibr">(Gregor and Hevner 2013)</ref>. VoC listening is about tapping into the "wisdom of the crowds"-the notion that the aggregation of information from external groups can garner better insights <ref type="bibr">(Surowiecki 2005)</ref>. This idea has been used to fuel "active" participatory approaches such as crowdsourcing, where organizations engage crowds via contests or other incentivized sharing structures. It has also been used as part of "passive" strategies such as opportunistic crowdsensing or big data analytics applied to crowd-generated data for social intelligence <ref type="bibr">(Zeng et al. 2010</ref><ref type="bibr">, Brynjolfsson et al. 2016)</ref>. Crowds can perform certain types of tasks fairly effectively, including cognitive tasks such as whether a given product will be successful or whether a product has issues <ref type="bibr">(Surowiecki 2005</ref><ref type="bibr">, Sunstein 2006</ref>). An important consideration is the amount of task-related information available to the crowd-wise crowds are able to leverage their knowledge, experiences, and intuition <ref type="bibr">(Sunstein 2006)</ref>. Additionally, crowd wisdom also embodies the following characteristics: diversity of opinion, independence, decentralization, and suitable aggregation mechanisms <ref type="bibr">(Surowiecki 2005)</ref>. These characteristics of wise crowds highlight  <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> three important implications for VoC listening platforms: (1) impact of tasks and objectives, (2) attributes of VoC channels, and (3) robustness of signal detection methods. Effective VoC listening platforms must carefully consider the design implications of each of these.</p><p>1. Impact of Tasks and Objectives: The major value proposition of crowd wisdom is that it can facilitate more accurate, timelier insights, leading to better outcomes <ref type="bibr">(Sunstein 2006</ref><ref type="bibr">, White et al. 2013)</ref>. Utility can be a relative concept, closely related to what the insights are being used for and by whom. For instance, the PredictIt political stock market's predictions about election outcomes might be used by investors speculating in prediction markets, journalists covering the election, candidate supporters eager to know who might win, and special interest groups looking to get a jump on prospective winners <ref type="bibr">(Surowiecki 2005)</ref>. VoC listening platform design needs to take into account stakeholder trade-offs based on respective risk tendencies, operational constraints/ capabilities, and major objectives.</p><p>2. Attributes of VoC Channels: The diversity, independence, and decentralization of user content generated in VoC channels may have important implications for crowd wisdom-gathering capabilities <ref type="bibr">(Surowiecki 2005)</ref>. Channels such as search queries and various social media platforms vary in quality, recency, uniqueness, frequency, and salience of content created <ref type="bibr">(Abbasi and Adjeroh 2014)</ref>. They also encompass varying social network structures that can impact information diffusion patterns <ref type="bibr">(Kwak et al. 2010</ref>). Furthermore, these channels may differ in terms of usage intentions. For instance, search query volume primarily reflects information acquisition patterns <ref type="bibr">(White et al. 2013</ref><ref type="bibr">, Brynjolfsson et al. 2016)</ref>, whereas forums are used for acquisition, dissemination, and sensemaking via conversations <ref type="bibr">(Abrahams et al. 2015)</ref>, and Twitter is commonly used for larger-scale broadcasting/ dissemination <ref type="bibr">(Kwak et al. 2010)</ref>. Consequently, VoC listening platform designers must understand crosschannel implications to "prioritize channels based on value" and "justify a more strategic investment" <ref type="bibr">(Davies 2015, p. 6)</ref>.</p><p>3. Robustness of Signal Detection Methods: Suitable aggregation mechanisms are essential for effectively leveraging crowd wisdom <ref type="bibr">(Surowiecki 2005)</ref>. These aggregation mechanisms must perform signal detection, the process of disentangling signal insights from noise <ref type="bibr">(Cassino 2016)</ref>. As <ref type="bibr">Abrahams et al. (2013, p. 871</ref>) note, detecting "whispers of useful information in a howling hurricane of noise" is a huge challenge, and filters are needed to extract meaning from the "blizzard of buzz." Inadequate signal detection methods can dramatically diminish the utility of VoC listening platforms, and effectively detecting signals from unstructured user-generated channels remains difficult <ref type="bibr">(Browne et al. 2015)</ref>.</p><p>On the basis of these important design implications, we propose a framework for examining VoC listening platforms (depicted in Figure <ref type="figure">1</ref>). Listening tasks and objectives are represented in the form of stakeholders, event types of interest for monitoring, and the importance of different event detection metrics such as accuracy and timeliness. VoC channels with varying characteristics are represented, including social Based on <ref type="bibr">Provost and Fawcett (2013)</ref> and <ref type="bibr">Blattberg et al. (2008)</ref>, the design decision process follows three steps: (1) attaining inputs from the stakeholder reflecting their trade-offs associated with model performance outcomes; (2) uncovering the interplay between design elements and model performance; (3) combining stakeholder inputs, model performance, and design elements to prescribe the best design choices that satisfy stakeholder priorities.</p><p>Two types of information constitute the necessary stakeholder inputs for the decision process. On the one hand, operations managers may have a different threshold for each metric based on their operational constraints and capabilities-for example, "our team cannot handle more than a certain signal volume, necessitating a higher precision threshold." We denote these minimum thresholds, which are similar to those in <ref type="bibr">Provost and Fawcett (2013)</ref>, as mP, mR, and mT. On the other hand, business and risk managers may have varying preferences for these metrics. For example, regulators need to take costly auditing actions for an adverse event, resulting in lower tolerance for false positives. By contrast, firms may be more likely to trade precision for better and timelier recall so that they can proactively cope with adverse events. We denote these preference weights as wP, wR, and wT. These weights are analogous to monetary costs and benefits, reflecting decision maker trade-offs. For example, a higher wP and a lower wR implies that the stakeholder associates a higher cost with false positives than false negatives.</p><p>Next, we use analysis of variance (ANOVA) and logit regression to uncover the impact of design elements on signal detection performance metrics Y ijklm {Precision, Recall, Timeliness}. The event type E k is a between factor (nested under the event type). The online user-generated channel D i , signal detection method M j , and temporal granularity T l are within factors. With S m standing for individual events, S/E m(k) denoting events within each type, and &#949; m(ijkl) as an error term, the structural model describing the sources of variance becomes</p><p>Accordingly, we can predict the possible modeling performance for each of the factorial (design element) combinations by calculating the marginal means, denoted as P ijklm , R ijklm , and T ijklm .</p><p>Finally, we generate the expected value of any design combination by considering preference weights and modeling performance <ref type="bibr">(Blattberg et al. 2008, Provost and</ref><ref type="bibr">Fawcett 2013)</ref>. The design element optimization is formulated as follows:</p><p>where T ijklm is a linearly transformed timeliness measure ranging between 0 and 1. Figure <ref type="figure">2</ref> illustrates an example of the design decision process <ref type="bibr">(Blattberg et al. 2008)</ref>. Stakeholders first provide design choices regarding relevant design elements (highlighted in green) specific to their contexts. In this example, the focused event type is product recalls; the processed data are from search queries and forums; the selected temporal granularities are daily, weekly, and monthly; and the chosen models are machine learning (ML) mention models and advanced models. Within this design choice solution space, signal detection performances for each of the design element combinations can be calculated. Combined with stakeholders' performance thresholds and metric weights, the platform design decision process can leverage ANOVA/logit regression and design Information Systems <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> element optimization to identify the optimal design choice leading to the best expected value, which is search query-advanced models-daily (path in solid arrows). At this point, the stakeholder can choose to accept the design, tweak its preference weights, or expand its design choices (e.g., consider alternative VoC channels) to continue to exploring designs with potentially better expected values. We later provide empirical results to demonstrate how the framework can serve as a decision aid for VoC listening platform design. The framework affords opportunities for examining the adverse event detection capabilities of different design configurations, as well as the overall impact of various platform design elements. In the ensuing section, we discuss each component of the framework, including the state of the art, limitations, and key gaps.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">VoC Listening: Related Work and</head><p>Research Gaps 3.1. VoC Listening Stakeholders and Types of Adverse Events VoC listening platforms are relevant to several stakeholders, including regulators, manufacturing firms, consumers and advocacy groups, and investors. In the United States, regulatory agencies such as the FDA, NHTSA, and Consumer Product Safety Commission are actively involved in postmarketing surveillance <ref type="bibr">(Chen et al. 2009</ref><ref type="bibr">, Abrahams et al. 2015</ref><ref type="bibr">, Sarker et al. 2015)</ref>. Manufacturing firms monitor their products to proactively mitigate risk, including "costs of managing reverse flow of products, disposal costs, restitution costs, and legal and liability costs due to any litigation" as well as indirect costs such as "loss of brand image and erosion of market value" <ref type="bibr">(Hora et al. 2011, p. 766)</ref>. Similarly, investors/markets are interested in detecting product issues that may impact their stock portfolio <ref type="bibr">(Chen et al. 2009)</ref>.</p><p>The key objectives of VoC listening for risk management are better and faster identification of adverse events <ref type="bibr">(Zabin et al. 2011</ref><ref type="bibr">, Yang et al. 2014)</ref>. From a big data analytics perspective, these objectives translate into three primary metrics: precision, recall, and timeliness. Precision and recall measure the ability to accurately identify adverse events. Recall denotes detection rate, whereas precision is a measure of falsepositive rate, with implications for "alert fatigue." Timeliness is how much earlier an adverse event can be detected, either in comparison with the point in time when the event transpires or relative to status quo detection methods <ref type="bibr">(Hora et al. 2011)</ref>. It is a significant VoC listening objective because earlier detection can expedite remedial actions, lessening social and monetary costs. Timely detection allowed Johnson &amp; Johnson to efficiently recall 31 million units of Tylenol in 1982, and Mattel was able to recall nearly 1 million toys containing lead-based paint in 2007. In both cases, early detection allowed product to be pulled from the supply chain before it adversely impacted consumers <ref type="bibr">(Hora et al. 2011)</ref>. Conversely, analysis of search query volume data could have allowed one to two years earlier detection of a dangerous adverse drug event causing hyperglycemia-potentially exposing one million fewer consumers to the event <ref type="bibr">(White et al. 2013)</ref>.</p><p>It is important to note that different stakeholders might define "better and faster" differently. For instance, a regulatory agency with a panoramic view of an entire industry encompassing thousands of products might have less bandwidth for false positives than a specific manufacturing firm with a single product line and more available monitoring resources. We underscore this point in our evaluation section by including results from the vantage point of regulators (industry level) and an individual firm. Additionally, we provide a platform design decision process component to consider the trade-offs facing different stakeholders and help them find the best design choices based on their thresholds and preferences, which is described in more detail in Section 3.4.</p><p>Past studies have examined adverse events that vary with respect to the product or nature of the event. For instance, some studies have analyzed events pertaining to specific categories of products such as pediatric or cancer medications (Hadzi-Puric and Grmusa 2012). Others have focused on events related to a set of manufacturers, such as Honda, Toyota, and Chevrolet <ref type="bibr">(Abrahams et al. 2015)</ref>. In the context of postmarketing drug listening, <ref type="bibr">Sarker et al. (2015)</ref> observe that most prior studies have examined a maximum of 5-10 products. Other studies have emphasized the importance of examining a wider set of products and event characteristics such as product recalls, safety communications, ongoing reviews, and severe warnings <ref type="bibr">(Hora et al. 2011, Abbasi and</ref><ref type="bibr">Adjeroh 2014)</ref>. In a broader review of social listening research spanning multiple industries and event types (including manufacturing defects, newspaper complaints, consumer electronic experiences), <ref type="bibr">Abrahams et al. (2013)</ref> also noted that most studies had relied on a single event type. From a VoC listening platform perspective, there remains a need to examine an array of important adverse event types related to multiple stakeholders.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">VoC Channels</head><p>Prior studies have noted the limitations of overreliance on any single data source, including uneven reporting patterns from consumers because of a lack of awareness of that particular channel <ref type="bibr">(Yang et al. 2014</ref><ref type="bibr">, Abrahams et al. 2015)</ref>. For instance, in the context of adverse drug events, <ref type="bibr">Xu and Wang (2014)</ref> find that FAERS yielded detection precision rates below 2.5%-meaning 39 out of every 40 alerts triggered was a false positive. This is consistent with our own evaluation results presented later in Section 6, in the context of pharmaceutical and automotive events. Inevitably, monitoring teams relying on such data sources must be conservative in their assessments as a result of fewer potential needles in the proverbial haystack.</p><p>Relevant alternative user-generated content channels are those encompassing consumer-contributed content <ref type="bibr">(Yang et al. 2014)</ref>. The most common categories incorporated in past studies are social media such as discussion forums and Twitter <ref type="bibr">(Abrahams et al. 2015</ref><ref type="bibr">, Lardon et al. 2015</ref><ref type="bibr">, Sarker et al. 2015)</ref> and search query logs <ref type="bibr">(White et al. 2013)</ref>. Twitter test bed sizes have ranged from a few thousand tweets containing a specific product name to billions of tweets mentioning an entire category of products (e.g., "cancer drugs") <ref type="bibr">(Sarker et al. 2015)</ref>. Discussion forums utilized were primarily consumer or product specificfor instance, the Honda-Tech forum for Honda issues <ref type="bibr">(Abrahams et al. 2013</ref>) and health discussion forums such as MedHelp, Drugs.com, and DailyStrength for adverse drug events <ref type="bibr">(Yang et al. 2014</ref><ref type="bibr">, Sarker et al. 2015)</ref>. Query log data have typically been attained from major search engines such as Google, Bing, or Yahoo!, and they usually include search query frequencies over time <ref type="bibr">(Karimi et al. 2015)</ref>.</p><p>Prior studies have typically focused on a single channel. However, these channels exhibit different characteristics with respect to credibility, frequency, and salience <ref type="bibr">(Agarwal et al. 2010, Abbasi and</ref><ref type="bibr">Adjeroh 2014)</ref>. For instance, on the one hand, social media channels such as Twitter and certain health forums are prone to spam, resulting in lower credibility <ref type="bibr">(Karimi et al. 2015)</ref>. On the other hand, forums have lower volume of content than Twitter and search queries but exhibit greater salience-forum postings are capable of incorporating greater background and context than a 140-character tweet and far more relative to a query encompassing a few search terms <ref type="bibr">(Abbasi and Adjeroh 2014)</ref>. Examining the user journey across multiple channels has become a major area of research with applications in e-commerce and marketing (e.g., customer journey and path-to-purchase) <ref type="bibr">(Song et al. 2014)</ref>. Similarly, there is a need to examine the effectiveness of different channels in the context of VoC listening. However, it remains unclear what the trade-offs of the user-generated content channels are with respect to detection rates, false positives, and timeliness of signals <ref type="bibr">(Sarker et al. 2015)</ref>. From a VoC listening platform perspective, the lack of cross-channel studies indicates a paucity of insights that can guide multichannel listening strategies.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3.">Signal Detection Modeling Methods</head><p>Signal detection methods for identifying adverse events can be broadly grouped into two closely related categories: basic mention models and machinelearning-based mention models <ref type="bibr">(Abrahams et al. 2013</ref><ref type="bibr">, Karimi et al. 2015)</ref>. Basic mention models consider disproportionality in the occurrence of key product and incident-related tuples relative to overall occurrences of these terms. For instance, in the context of adverse drug event detection, basic mention models typically measure a combination of a drug reference, some reaction-related terms, and, in some cases, anatomy or drug administration-related terms <ref type="bibr">(Adjeroh et al. 2014</ref>). An example of this would be, "I have been taking Drug x and began experiencing headaches and pain in my lower back." Several basic mention models have been proposed that leverage these product-incident co-occurrence values as input. Here, we briefly describe these methods; further details appear in Online Appendix A.</p><p>Relative risk (RR) is computed as the relative ratio of observed and baseline counts p(j|i)/p(j), where i and j denote mentions of product and effect, respectively (DuMouchel 1999). One noted limitation of RR is its susceptibility to sampling variability in situations where the observed and baseline counts are both small <ref type="bibr">(Karimi et al. 2015)</ref>. Proportional reporting ratios (PRR) extends RR by considering the cooccurrence of i and j relative to the occurrence of j in instances without i <ref type="bibr">(Yang et al. 2014)</ref>. Given that i &#8242; represents all product i instances devoid of j, PRR can be interpreted as p(j|i)/p(j|i &#8242; ). Another commonly used mention model is the reporting odds ratio (ROR) (Hadzi-Puric and Grmusa 2012): p(j|i)/p(j &#8242; |i)/p(j|i &#8242; )/ p(j &#8242; |i). In benchmarking studies on two data sets encompassing adverse drug event reporting system, medication order, and abnormal laboratory result instances, ROR performed comparably in some circumstance, and slightly better in others, relative to comparison methods such as PRR <ref type="bibr">(Liu et al. 2013)</ref>. Information component (IC) is an information theorybased measure that leverages the mutual information between i and j. IC is used by the World Health Organization, often with better results than other methods <ref type="bibr">(Lindquist 2008)</ref>. It can be computed as log 2 (p(j|i)/p(j)). Basic mention models have also been used in non-health event detection contexts. <ref type="bibr">Abrahams et al. (2012)</ref> have developed a "smoke words" method in which terms were weighted based on their occurrence in different types of automotive adverse event mentions.</p><p>Other mention models have incorporated supervised or unsupervised machine learning methods that learn patterns involving product and incident terms derived from dictionaries, lexicons, and/or thesauri. For instance, <ref type="bibr">Yang et</ref>   <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> support vector machine (SVM) classifiers that included semantic features such as bag-of-words, product names, and incident term lexicons to determine whether a particular product reference was a valid mention. <ref type="bibr">Abrahams et al. (2013)</ref> train linear SVM and na&#239;ve Bayes classifiers coupled with a term list feature set to detect vehicle component mentions. Similarly, in the health context, Liu and Chen (2013) develop an SVM classifier that used a custom kernel for relation extraction. For all sentences containing drug and reaction keywords, they derived the shortest dependency path using the StanfordCoreNLP package. Next, they replaced all path tokens with higherfrequency class tokens encompassing part-of-speech tags and entities (e.g., drug, event). The custom kernel function K(x,y) was simply the number of common features between the modified shortest dependency path strings for sentences x and y. SMART combined a logistic regression classifier with a feature set including style, semantic, and product attributes to detect defect mentions in consumer electronics and vehicle discussion forums <ref type="bibr">(Abrahams et al. 2015)</ref>. <ref type="bibr">Sampathkumar et al. (2014)</ref> represented product name, relation keywords, and incidents as hidden states in a hidden Markov model (HMM). Their HMM allowed these three named entities to occur in any order within a message and also included a fourth state for all "other" words within the message. A popular unsupervised method has been association rule mining, where product-incident co-occurrence patterns are derived based on their support and confidence scores <ref type="bibr">(Yang et al. 2014)</ref>.</p><p>Given mention occurrence frequencies over time, time-series analysis can be performed at different temporal granularities (e.g., daily, weekly, monthly, yearly). Similar to prior temporal prediction studies (e.g., <ref type="bibr">Fang et al. 2013)</ref>, most signal detection methods use windowing to apply temporal association rules or z-score thresholds to the time series at each time period t i &#8712; T = {t 1 , . . ., t g }. This is done by computing these measures over the dynamic time window t 1 to t g or in some cases beginning with some "training period" to allow suitable thresholds for earlier time periods close to t 1 <ref type="bibr">(Jin et al. 2010</ref><ref type="bibr">, Yang et al. 2014)</ref>. Figure <ref type="figure">3</ref> illustrates how the thresholds &#964; g and &#964; g+1 are used for the time-series windows ending at t g and t g+1 , respectively. In this case, the signal at t g+1 triggers an alert with timeliness t et g+1 . To avoid future leaks, for instance, let us assume we are building a monthly model over data beginning in January 2008 (t 1 ), and we are currently at January 2009 (t g ) in our timeseries windowing. The model will only use all data from January 2008 through December 2008 to try to detect a signal. Furthermore, let us assume that this triggers a spike in June 2008. For timeliness purposes, the signal time period will still be considered January 2009 (t g ) because the signal was detected using data up to that point in time. Windowing is performed across all data in the test bed, until t f . In summary, existing mention models mostly consider co-occurrence between individual products and incidents. However, many adverse events pertain to product interactions, entire categories of products, and/or incidents encompassing multiple issues. Furthermore, they fail to weight different mention components based on their implications for precision, recall, or timeliness in diverse contexts. Not surprisingly, performance results have varied, with recall rates often below 50% <ref type="bibr">(Adjeroh et al. 2014)</ref>. Those that examine precision have observed that such methods are prone to high false-positive rates-in many cases, 75% of signals or higher <ref type="bibr">(Adjeroh et al. 2014)</ref>. It is unclear how effective existing mention models are when applied to various VoC channels, for a broad array of products and adverse event types. There is a need for more robust signal detection methods beyond basic "mention" models, capable of enhanced precision, recall, and timeliness.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4.">Considering Stakeholder Priorities</head><p>Stakeholders have varying priorities based on their risk tendencies, operational constraints/capabilities, and major objectives. These priorities can have a significant impact on the design elements of a VoC listening framework. The expected value framework <ref type="bibr">(Provost and Fawcett 2013)</ref> and profitability framework <ref type="bibr">(Blattberg et al. 2008</ref>) provide an analytical approach that can guide the design decision process. They begin by decomposing the focal problem into all the possible outcomes of implementing an analytics model, obtain values (e.g., costs and benefits) from stakeholders to factor in trade-offs associated with different outcomes; identify design elements affecting the occurrence probabilities of different outcomes; weight the values by the probabilities of reaching a consolidated expected value; and provide prescriptions about which design element combination and what level of modeling performance (performance thresholds) are needed to accomplish a desirable expected value.</p><p>Whereas both frameworks are well suited for guiding evaluations and providing prescriptions for designing VoC platforms, each focuses on binary classification in a business context, necessitating adaptation to our signal detection context for a few reasons. First, for binary classification problems, the confusion matrix determining the possible outcomes is readily attainable. For our signal detection context, however, the possible outcomes worth consideration are beyond the confusion matrix. For instance, for a given event, we only need a single true positive. Additionally, the negatives in signal detection contexts are often harder to understand-an unknown unknown. Second, the temporal aspect of signal detection is critical in practice, which is not well captured in a binary classification context. Finally, in adverse event detection contexts, costs and benefits generally require more time and effort to obtain (e.g., hard to convert the societal benefits and costs into monetary values). Therefore, it is unclear how effectively existing analytic approaches for guiding design decisions can be adapted to the VoC listening platform context.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.5.">Inclusion of Sentiment Information and</head><p>Signal Fusion Sentiment analysis has seen limited usage in past studies examining adverse events in VoC channels. <ref type="bibr">Sarker and Gonzalez (2015)</ref> include sentiment polarity scores derived using the popular SentiWordNet lexicon <ref type="bibr">(Esuli and Sebastiani 2006)</ref>. They include the overall negative sentiment polarity as an input feature in their SVM classifier and find that the inclusion of sentiment provided a small lift in mention detection accuracy on their social media data set. Similarly, <ref type="bibr">Yang et al. (2013)</ref> include an affect lexicon. <ref type="bibr">Abrahams et al. (2012)</ref> use the Harvard General Inquirer dictionary of positive and negative keywords and find that it did not improve vehicle defect identification from online discussion forums. They conclude that general-purpose lexicons might be insufficient to capture nuanced opinion cues appearing in domain-specific online forums. In the context of search, <ref type="bibr">Turney and Littman (2003)</ref> propose a simple yet effective method pointwise mutual information method for deriving the sentiment of a search term: by comparing the search query volume for the term plus a set of positively oriented words (e.g., good, positive) and the search query volume for the term and a set of semantically opposed words (e.g., bad, negative).</p><p>Similar to sentiment analysis, signal fusion methods have seemingly limited usage for identifying adverse events despite potential for enhancing precision and recall by combining results across channel-specific signals via fusion schemes that are analogous to ensemble voting methods used in meta-learning <ref type="bibr">(Adjeroh et al. 2014)</ref>. Given that certain user-generated content channels such as Twitter and search have limited salience <ref type="bibr">(Abbasi and Adjeroh 2014)</ref>, inclusion of sentiment information could provide an important context refinement regarding user intention in these channels <ref type="bibr">(Sharif et al. 2014</ref>). In the same vein, the potential for signal fusion methods to enhance VoC listening capabilities remains underexplored.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Mention Model and Genetic Algorithm-Based Signal Detection</head><p>Robust signal detection is essential for effectively tapping into crowd wisdom <ref type="bibr">(Surowiecki 2005)</ref>. Here, we describe the basic mention model and then discuss our proposed novel heuristic-based method. We use examples related to health adverse drug events, but the methods are generalizable to an array of adverse product event contexts. We later evaluate them on health and automotive test beds.</p><p>To identify potential incident references, brand/ product, product attribute, and consumer experience lexicons are utilized. An automated tagging tool was developed to assign lexicon tags to references appearing in VoC channel documents. In the health adverse drug event context, these lexicons include drug, anatomy, reaction, and drug administration keywords. For example, the statement "I've experienced chest pains ever since I started taking Chantix" would be tagged as "I've experienced &lt;ANATOMY&gt; &lt;REACTION&gt; ever since I started taking &lt;DRUG&gt;." For word-sense disambiguation, we use the CMU part-of-speech tagger designed specifically for short <ref type="bibr">Abbasi et al.:</ref> User-Generated Signals for Adverse Event Warnings Information Systems <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> informal texts to help improve the likelihood that anatomy, side effect, and administration tags were applied appropriately.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.">Mention Model for Signal Detection</head><p>The basic mention model incorporated in this study can be described as follows. For each product E in our database, we build a fully unsupervised time series. Let t k &#8712; T x = {t 1 , . . ., t g } signify a given time window, where t g is the current time period of the analysis and t g is less than the final time period t f . Let C Dn represent the number of product names (D) associated with E that appear in a document n. Let {d 1 , . . ., d N } signify the set of documents occurring during t k within a given channel, where each C Dn &#8805; 1. Furthermore, in our health context example, let C An , C Rn , and C Mn represent the accompanying product attribute and customer experience lexicons. These would be number of anatomy (A), reaction (R), and administration (M) terms present in document n, respectively. The aggregated raw score for time t k is then computed as</p><p>where &#181; g and &#963; g are the mean and standard deviation, respectively, across all t in T x plus the training period (see Figure <ref type="figure">3</ref>) where s(t k ) &gt; 0. For a given event time series, the basic model considers an alert at time t k if z(t k ) &gt; &#964; g , where &#964; g is a threshold for the current window. If t k is less than the event time period t e , it is considered a positive signal with timeliness t et g ; T x can vary depending on the resolution of the signals-such as daily or monthly time models, as well as the value of the current window time period t g .</p><p>It is important to note that a single, fixed set of anatomy, administration, and reaction terms are utilized for all drugs. Each product E is represented as a single time series where the y-values are the ztransformed s(t k ). Each event is a spike that exceeds the z-score threshold. Hence, the method is purely unsupervised, without use of any event knowledge a priori. To ensure avoidance of future leaks, neither the drug, reaction, anatomy, and administration terms nor the spikes that are generated use any event information.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2.">Genetic Algorithm-Based Signal Detection</head><p>Effective signal detection entails disentangling signal from noise <ref type="bibr">(Sunstein 2006)</ref>. One of the biggest limitations of prior mention models has been that they adopt a "one-size-fits-all" approach-applying the same features, weights, and statistical patterns to a diverse set of products, channels, and user experiences. Basic models apply a cookie-cutter disproportionality idea to an array of product-incident mention tuples, resulting in low precision rates. Conversely, supervised machine learning methods offer better precision but often lack generalizability necessary to garner adequate recall (because of limited diversity of the training data with respect to channels and products). Our proposed genetic algorithm-based signal detection (GASD) method attempts to address these concerns by building signals capable of better accounting for product-and channel-specific characteristics. The two key aspects of the method are its (1) objective function, which rewards the creation of signals that garner fewer, potentially higher-quality, alerts faster; and (2) the weighting method, which allows better contextualization of references to product, attribute, and user experience terms for each individual product. GASD attempts to better harness the diversity of wise crowds for enhanced aggregation, in an unsupervised manner devoid of overfitting. The details are as follows.</p><p>GASD learns time-series-specific weights for various product, incident, and experience terms. Extending our drug example, let F Dn represent the occurrence vector of drug terms in document n for product E, and let W D denote the vector of weights for drug terms where each W Dx &#8712; {0, 1/(b -1), 2/(b -1), . . .,1}, and b indicates the number of discrete weight intervals. In GASD, s(t k )</p><p>, with the objective of finding suitable values for W D , W A , W R , and W M . Within a population of solutions P, we represent each solution p q as a binary "bit" string encompassing values for all four sets of terms. Each weight value in p q is represented using h bits such that there are 2 h = b possible weight values for each term. For each p q , the fitness function f(p q ) is used to evaluate each signal s(t k ) within and across each window T x . The fitness function considers the timeliness of the signal, the importance of incorporating key reference terms, and provisions to alleviate false positives:</p><p>where t g denotes the end of a given window T x , t k &#8712; T x indicates one of the time periods that triggers an alert, and D(s(t k )), R(s(t k )), and so forth, indicate the number of drug, reaction, etc., terms appearing in the top r ranked list in period k based on WF values. The variable l denotes the total number of alerts triggered in T x , used to penalize the fitness value for signals generating excessive alerts. Further details regarding the GASD fitness function and weighting mechanism appear in Online Appendix H. Figure <ref type="figure">4</ref> shows the GASD formulation using the aforementioned fitness function and bit-string encoding. GASD is run for each sliding window instance t 1 to t g as previously illustrated in Figure <ref type="figure">3</ref>. For each subsequent generation, the selection probability of a solution is proportional to its f(p q ). Within the new solution set O, crossover is applied on adjacent solutions o q and o q+1 with probability c, and mutation is applied on individual bits within each p q with probability m. In results reported, c = 0.7, m = 0.001, and r = 20 were used (i.e., no tuning was performed, to avoid overfitting). Stopping criterion for genetic algorithms are an important consideration <ref type="bibr">(Aytug and Koehler 1996)</ref>. In our analysis, we observed that GASD consistently converged within 200 iterations; however, because run times were not a concern, we used a fixed 500-iteration stopping criterion (i.e., terminate after 500 generations).</p><p>Figure <ref type="figure">5</ref> presents a short illustrative example of the effectiveness of GASD. The chart on the left depicts the GASD signal for the drug Actos, relative to the mention model (right). The FDA announced an investigation on 9/17/2010 for bladder cancer in patients using Actos over an extended period (denoted with an X). The horizontal lines indicate an alert threshold, and TP and FP denote true/false alerts. From the figure it is evident that although both signals appear similar, GASD's term weighting is able to allow earlier detection and fewer false positives by dynamically weighting various drug, reaction, anatomy, and administration keywords, resulting in subtle yet impactful changes in signal strengths.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Evaluation Test Bed and Design</head><p>Two adverse product event test beds were incorporated. In the main paper, we report results from the health industry, related to adverse drug events. Online Appendix C presents results from the automotive Information Systems <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> industry for adverse automotive events. For our health test bed, we collected data on all drugs that had a first-time FDA drug alert for adverse drug events between 2011 and 2013, resulting in 143 events related to 133 unique drugs associated with a myriad of illnesses and ailments, including diabetes, blood pressure, cholesterol, cancer, depression, chronic pain, birth control, insomnia, Parkinson's, arthritis, seizures, etc. The events corresponded to four types: drug safety communications are first notifications of a new adverse reaction problem; ongoing reviews indicate that the FDA is investigating whether there is a problem; FDA news is often forwarded from pharmaceutical company self-reported issues; and product recalls are typically due to a manufacturing, packaging, or labeling issue, as opposed to a drug issue. Data from three user-generated content channels were collected: Twitter, forums, and search logs. Table <ref type="table">1</ref> presents an overview. Approximately 12 million tweets containing drug-name keywords spanning 2006 to 2014 were gathered through Topsy's API. Over 5 million postings from 10 popular health forums were attained using web crawlers. The postings spanned the time period 2000 onward. These messages were converted to sentence chunks, resulting in 26 million forum instances in the test bed. This was done because the forum messages were lengthier and often contained discussion of multiple topics. Sentence-level analysis resulted in better performance and information units that were more focused and consistent with the tweets and search queries. Search query frequencies over time were attained from publicly available online sources, as done in prior studies <ref type="bibr">(Brynjolfsson et al. 2016)</ref>. In particular, we used Google Trends to attain search query volume over time at different temporal granularities for terms in our drug, reaction, anatomy, and administration lexicons, as well as search term co-occurrence volumes. In addition, 6.2 million reports submitted to FAERS were also incorporated in the baseline evaluation to illustrate the value of search and social channels. Consistent with prior studies, for each report, the set of drugs and reaction terms were used to build the FAERS signals <ref type="bibr">(Xu and Wang 2014)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.">Baseline Evaluation: Comparing Existing Mention Models and Examining Regulatory Databases</head><p>Before investigating our core research questions, we conducted two baseline evaluations. The purpose of the first was to illustrate that the baseline mention model utilized in this study was indeed indicative of the types of performance results attained using methods from prior studies. We incorporated several representative comparison methods discussed in the literature review, including basic mention models, machine-learning methods used in prior adverse product event detection studies, and general event detection methods. Performance on the Twitter and forum test beds was examined at three temporal resolutions for the signal time series (day, week, and month). For both channels, we only used data from 2008 onward to allow for better comparison of acrosschannel performance. For each method, a mention frequency time series was constructed with &#964; tuned using a grid search. A step interval of 1 was used to compute mean and standard deviation (the basis for the z-score threshold) over a growing window T.</p><p>For each event, windowing was performed until the FDA first report date. Consistent with prior work, all methods were evaluated using the standard aforementioned metrics: recall, precision, and timeliness. However, given that our task entails identifying adverse events earlier than the first official mention, "positives" were only those signals that occurred prior to the regulator first-report date for that particular drug event. For each such identified "positive" signal, a determination of true positive (TP) or false positive (FP) was made using a two-stage approach. First, the key drug, reaction, and anatomy keywords appearing in the signal were automatically compared against those appearing in the regulator descriptions. If the similarity was below a certain threshold, the signal was automatically rejected as a false positive. For those above a threshold, an independent domain expert examined a sample of documents pertaining to the signal (e.g., the underlying tweets, postings, queries) to determine relevance. Precision and recall were computed as TP/(TP + FP) and (earlier detected events)/(total events), respectively. Timeliness was derived as the earliest average number of days for a detected event, relative to the first FDA event date. Details regarding the comparison mention models and the evaluation results appear in Online Appendix B. Here, we summarize results for the mention model (labeled "Mention" in Figure <ref type="figure">6</ref>) and the average results for the three comparison categories of methods evaluated: basic co-occurrence mention models (Basic), machine-learning methods (ML) used in prior adverse product event studies, and general machine-learning methods used in event detection studies (General ML). "Mention" yielded comparable results to those utilized in prior research; it had the highest recall and precision values near the top as well. It also yielded timelier results than other basic co-occurrence models. Overall, the results underscore some of the limitations of existing mention models alluded to in the related work section-they generated low precision (mostly below 20%) and, with the exception of certain daily models, also yielded recall rates below 50%.</p><p>As previously alluded to, prior studies have noted the limitations of spontaneous reporting databases such as FAERS <ref type="bibr">(Xu and</ref><ref type="bibr">Wang 2014, Yang et al. 2014</ref>). To illustrate the potential of alternative VoC channels, we ran the four baseline mention models (RR, PRR, ROR, and IC) on the 6 million FAERS reports in our test bed (mentioned in Table <ref type="table">1</ref>). Because the FAERS data set did not include text descriptions, only drug and side effect sets, we could not utilize our machine learning methods. We compared the precision, recall, and timeliness of FAERS versus Twitter and forums for the 143 events using these four mention models and found that Twitter and forums yielded significantly better performance across all three performance metrics, for all four mention models (all p-values &lt; 0.001). Figure <ref type="figure">7</ref> presents the daily, weekly, and monthly precision, recall, and timeliness results for FAERS, Twitter, and forums, averaged across the four baseline mention models. Consistent with prior studies (e.g., <ref type="bibr">Xu and Wang 2014)</ref>, FAERS garnered precision rates below 3% and recall rates that were 15-20 points lower than the social media channels. As expected, the timeliness of FAERS true-positive alerts was typically within three to five months of the official first notification date. This is not surprising because FAERS is the primary data source for many of those first notification dates in the first place. Conversely, the social media channels were much timelier.</p><p>Our baseline evaluation highlights the potential of alternative VoC channels relative to existing reporting databases, and it shows that the mention model incorporated in this study is representative of prior baseline models. In the following section, we incorporate this mention model as well as the proposed heuristic-based model to examine the effectiveness of user-generated content channels using more robust Information Systems <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> detection methods, and we assess the interplay between channels, event types, and modeling methods for VoC listening platforms.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.">Evaluation Results</head><p>To investigate our two research questions, we used a factorial design encompassing all three channels (Twitter, forums, and search), two types of signal detection models (basic mention model and GASD model), and three temporal granularities for the signal time series (day, week, and month). Initially, sentiment was excluded from the analysis. This resulted in 18 total combinations of signal detection methods. Once again, we only used data from 2008 onward for each channel to allow better comparison of across-channel performance. In all evaluations, &#964; = 1 was used for GASD as a default (i.e., no tuning), whereas, once again, the best setting for all comparison methods over a grid search was adopted.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.1.">Event Detection Performance</head><p>Table <ref type="table">2</ref> presents the results. GASD outperformed the mention model by a wide margin in terms of recall and precision on all settings (typically more than 10 points better). It also yielded timelier signals on most settings, generally detecting events 100-200 days earlier, or more. The recall rates for GASD were in the 55%-80% range, and it attained markedly better precisions rates than prior studies-above 34% for all settings and, in some cases, upwards of 50% or 60%. With respect to channels, forums and Twitter garnered higher recall than search (20 points better on average), but search attained precision rates that were typically at least 10-15 points better. Similarly, the classic trade-off between recall and precision was also observed with respect to temporal resolutions: daily models yielded the best recall but were also prone to the most false positives (likely as a result of greater volatility and noise). Recall rates across event  types were generally highest for ongoing reviews and drug safety communications. Not surprisingly, results on product recalls were lower because many of these events are devoid of any explicit reaction or anatomy terms (e.g., "pills chipped or broken in the packaging plant"), making it difficult to detect such signals via VoC channels.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.2.">ANOVA and Logit Regression to Examine</head><p>Interplay Between Event Types, Channels, and Models To examine the effect of event types, channels, and model types on our three performance metrics within the 18 model settings, we conducted an ANOVA for each of the 143 events. Specifically, we employed a three-way mixed ANOVA design (i.e., split-plot design). The event type was a between factor (nested under the event type). The channels and models were within factors, indicating that for the same event, signals were repeatedly extracted from different channel and model combinations. We used the weekly granularity models for all analyses to manage the complexity of introducing a fourth variable. Because recall in this setup was binary (i.e., either the event was detected or it was not), we used a logistic mixed model to analyze recall but with the same factorial structure. With S standing for the 143 individual events, A denoting the event type, B signifying channels, and C representing model types, and with S/A denoting events within each type, the structural model describing the sources of variance becomes</p><p>We do not present the results for precision here because of space constraints; however, the betweenfactor event type (A) and the within-factors channels (B) and models (C) all had a significant effect on precision (p-values &lt; 0.01). Table <ref type="table">3</ref> depicts the ANOVA results on timeliness. The main effects of event types (A), channels (B), and models (C) were all significant (p-values &lt; 0.01). Interestingly, there was a significant interaction effect between model and channel, as shown in Figure <ref type="figure">8</ref>(a). Among the three channels, the GASD model obtained the least gain in the search channel and the most gain in the forum channel. Finally, signals detected from the search and forum channels had better timeliness than those detected from Twitter. This latter result has interesting implications that we elaborate on later in Section 8. As depicted in Table <ref type="table">4</ref>, similar to precision and timeliness, event types (A), channels (B), and models (C) all had a significant impact on recall (p &lt; 0.01). Additionally, there was a significant interaction effect between event type and channel, as well as channel to model, and a significant three-way interaction among them. Figure <ref type="figure">8</ref>(b) illustrates the former interaction effect with the channel factor being the comparison basis. As previously alluded to, in general, forums and Twitter were better than search in terms of recall rates. However, the relative strength between forums and Twitter varied depending on event types, with forums performing better on product recalls, whereas the opposite was observed for ongoing reviews.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.3.">Impact of Including Sentiment and Signal</head><p>Fusion in Mention Model and GASD To examine the impact of sentiment, we trained a machine learning sentiment analysis classifier and ran it on the forum and Twitter data sets to derive  <ref type="bibr">, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> message-level sentiment polarity scores <ref type="bibr">(Hassan et al. 2013</ref><ref type="bibr">, Sharif et al. 2014</ref><ref type="bibr">, Zimbra et al. 2018)</ref>. For the search channel, the semantic orientation method proposed by <ref type="bibr">Turney and Littman (2003)</ref> and previously described in Section 3 was adopted. Details about the sentiment analysis methods utilized appear in Online Appendix D, along with the full evaluation results. Here, we summarize the key takeaways.</p><p>Figure <ref type="figure">9</ref> shows the precision and recall performance differences for the 18 models with negative sentiment polarity included in the models, relative to the 18 with no sentiment (presented in Table <ref type="table">2</ref>).</p><p>Positive values indicate higher performance with sentiment. Looking at the results, the models with sentiment tended to attain higher precision across the board. However, often the sentiment models also resulted in lower recall. The differences in precision and recall were especially pronounced on the discussion forums, where inclusion of sentiment resulted in marked improvements in precision and decreases in recall. Interestingly, forums are the only channel incorporated that is primarily for the discussion of product issues and experiences. Hence, in this channel, co-occurrence mentions devoid of sentiment information are more likely to result in false positives <ref type="bibr">(Sarker and Gonzalez 2015)</ref>.</p><p>GASD method from a manufacturer's vantage point, we present a brief case study from the perspective of the risk management group at Pfizer. We analyzed 20 products from their portfolio, some of which had adverse events that transpired during the time period between 2011 and 2013 (note that some drugs had no events). The events were associated with two types: drug safety communications and product recalls. We ran the mention and GASD models on all Twitter, search, and forum channel data in our test bed and computed precision, recall, and timeliness.</p><p>Table <ref type="table">5</ref> shows the evaluation results. GASD was able to detect 76%-84% of the events three to four years earlier. Risk management groups at such firms are often willing to have slightly lower precision (i.e., 33%-60%) for better, timelier recall. Given the size of their monitoring team, and the relatively fewer products needing monitoring compared with a regulatory agency, such groups are well suited to investigate one or two false alerts for each true positive-a far better ratio than the mention model. Although not depicted, GASD again had lower standard deviations on timeliness.</p><p>Figure <ref type="figure">11</ref> illustrates the value of the enhanced precision, recall, and timeliness enabled by GASD relative to mention models. Depicted are two Pfizer drugs for which an FDA event transpired. For each event, all GASD and mention model alerts are displayed (true and false positives). For example, an adverse event for the drug Revatio was first detected by GASD 22 months prior to the FDA announcement. In total, GASD had four true positive and two false positive alerts for this product. The figure highlights how signal quality with respect to precision, recall, and Information Systems <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> timeliness can impact practical value in real-world settings. Relative to GASD, not only does the mention model fail to detect the Revatio event but also it detects the Zithromax event 18 months after GASD. Furthermore, it generates more false-positive signals, which can cause "alert fatigue" over time, impacting the perceived usefulness of VoC listening capabilities.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.5.">Design Decision Process for</head><p>Different Stakeholders As alluded to, the proposed framework can also help different stakeholders to identify the best VoC listening platform design element combination based on their specific thresholds and preferences. Table <ref type="table">6</ref> presents an illustration from the health test bed involving the FDA and Pfizer. In our example, although both stakeholders have similar threshold requirements (i.e., minimum precision, recall, and timeliness), they have differing preference weights for performance metrics. As a regulator, the FDA may have more stringent requirements for precision because of the hefty cost of investigating many signals. Conversely, Pfizer may place greater weight on recall and timeliness to proactively address as many adverse events as possible for monetary and risk mitigation reasons.</p><p>We ran ANOVA and logit regression on our two stakeholder's event data sets, respectively, and we estimated marginal means for all the design element combinations (as described in Section 2). The results are depicted in Table <ref type="table">6</ref>. We find that for most (but not all) event types, the FDA should use GASD at weekly or monthly temporal granularities on different channels. By contrast, for the same event types, the best alternative for Pfizer is GASD running on the forum channel at daily intervals. The example indicates how our framework can take stakeholder inputs to determine the best design elements for signal detection. Depending on stakeholder constraints, some event types (e.g., FDA ongoing news in the fourth row) may not be viable for listening under the current thresholds. The example illustrates how our proposed design decision process can help stakeholders identify design elements based on their specific requirements.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="8.">Results Discussion and Conclusions</head><p>Collectively, the evaluation results demonstrate that event types, channels, and models heavily impact event signal detection performance. Table <ref type="table">7</ref> summarizes some of the key results pertaining to our two research questions. In general, GASD provided better precision, recall, and timeliness than the mention model. GASD's false-positive rates were two to five times lower than the mention model. The earlier detection rates for many events, across models and user-generated channels (i.e., often one to three years earlier), is also an interesting result that is consistent with some recent studies <ref type="bibr">(White et al. 2013</ref><ref type="bibr">, Adjeroh et al. 2014)</ref>. Performance was best on event types with greater salience. In the health test bed, examples included reviews and safety communication events, relative to product recalls. Social media channels provided higher recall but lower precision than did search. Within social media, forums and Twitter each performed better on certain event types (e.g., forums attained higher recall for product recall events in both the health and automotive test beds). Forums and search also yielded timelier detection than did Twitter.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="8.1.">Contributions to Research</head><p>We contribute to the emerging IS research on application of big data analytics to problems with societal implications (e.g., <ref type="bibr">Chen et al. 2012</ref><ref type="bibr">, Bardhan et al. 2015</ref><ref type="bibr">, Abbasi et al. 2016</ref><ref type="bibr">, Brynjolfsson et al. 2016)</ref>.</p><p>From a design science perspective, our contributions include our analysis framework and the GASD User-generated content channels can enable timelier detection of events for 50%-80% of events examined. Using search/social channels, events can be detected one to three years earlier than first reports in official event databases used by regulators and manufacturing firms. Whereas existing mention models suffer from low precision rates (i.e., &lt; 20% on most channels) when applied to search/social, the proposed GASD method garnered false-positive rates that were two to five times lower than those of the mention models, with precision as high as 50%. The inclusion of sentiment in the models enhanced precision, particularly in the online forums. Given that much of the discussion in these forums encompasses incidentrelated keywords, the extent of negative sentiment may constitute an important filtering mechanism. Similarly, signal fusion can further enhance detection recall and overall f-measures. The aforementioned results with respect to recall (i.e., detection rates), timeliness, and precision (i.e., percentage non-false alarms) were consistent across daily, weekly, and monthly models in test beds spanning two different industry contexts. What are the relevant interactions between channels, event types, and modeling methods, and what are their implications for the design of VoC listening platforms?</p><p>With respect to event types, performance was best on events where users can explicitly mention adverse experiences. In the health test bed, examples include ongoing reviews and drug safety communications. Conversely, drug recalls, which often stem from manufacturing, packaging, or labeling issues, were generally more challenging to detect with search/social channels. Social media channels examined (i.e., forums and Twitter) garnered higher recall but lower precision relative to search. Hence, social media has greater signal and noise, possibly because of the competing effects of greater salience and contextualization on one hand and credibility implications on the other (e.g., spam). Forums and search were timelier than Twitter. This is counterintuitive with findings in other domains such as financial services. The IS literature on crowd-generated data, user motivations for sharing, and online privacy may offer alternative explanations for this effect. Between the two social media channels, forums were better at detecting product recall events in the health and automotive test beds, whereas Twitter provided better detection capabilities for ongoing reviews in the health test bed. The proposed framework can prescribe the best VoC listening platform design choices for a given set of inputs reflecting the stakeholder's performance trade-offs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Abbasi et al.: User-Generated Signals for Adverse Event Warnings</head><p>Information Systems <ref type="bibr">Research, 2019</ref><ref type="bibr">, vol. 30, no. 3, pp. 1007</ref><ref type="bibr">-1028</ref><ref type="bibr">, &#169; 2019 The Author(s)</ref> method. The existing literature on adverse event detection has been disparate and fragmented. By using the crowd-generated data literature as a kernel theory, our framework provides a holistic lens for examining key considerations related to VoC listening. An extensive evaluation across two large test beds demonstrated the utility and robustness of the insights generated by our framework and of the GASD method versus existing basic, machine learning, and general-purpose methods. Moreover, the framework can be used to prescribe VoC listening platform design configurations for a given set of stakeholder inputs related to performance trade-offs and constraints. IS scholars have framed the differences between studies examining prediction versus explanation and the importance of both <ref type="bibr">(Shmueli and Koppius 2011</ref><ref type="bibr">, Agarwal and Dhar 2014</ref><ref type="bibr">, Bardhan et al. 2015)</ref>. Our analysis framework and the empirical insights produced also provide opportunities for future theory development that offers rich explanations. Using <ref type="bibr">Gregor and Hevner's (2013)</ref> classification, our work makes contributions to "nascent theory" that is not yet fully emerged. Three examples are as follows.</p><p>1. Timeliness and Channel Usage Motivations: Our findings regarding the timeliness of search and forums relative to Twitter are interesting because prior studies on topics such as the relation between social media and stock performance found Twitter to be a stronger lead indicator than forums <ref type="bibr">(Das and</ref><ref type="bibr">Chen 2007, Bollen et al. 2011)</ref>. As noted in Section 2, the crowd-generated data literature suggest that channel usage intentions (e.g., information acquisition, discussion, and dissemination) might explain the timeliness differences. From a "customer knowledge acquisition journey" perspective, it is possible that internet users sequentially use search to acquire, forums to discuss, and Twitter to disseminate. Alternatively, the earlier use of search and forums, at least in the health test bed, could be attributable to the relatively sensitive nature of the health domain <ref type="bibr">(Anderson and Agarwal 2011)</ref>. Search constitutes a more private information-gathering option. However, a similar effect was observed in the automotive test bed, a seemingly less sensitive domain. The sharing literature notes that peoples' motives for sharing information and insights may be driven by a desire to help others <ref type="bibr">(Fichman et al. 2011)</ref>, or for social capital <ref type="bibr">(Wasko and Faraj 2005)</ref>, and forums provide a more conducive channel for sharing such information compared with Twitter. Our findings suggest an opportunity for studies exploring how certain segments of the population prefer to discuss or disclose their adverse experiences, over time.</p><p>2. Recall and Salience of Different Event Types and Channels: Our product recall event type had low detection rates in the health test bed. However, in the automotive context, these types of events had the highest detection rate, likely because of the abundance of sensory and diagnostic cues. For instance, customer mentions included references to sounds, smells, vibrations, and warning lights that are more salient and easily connectable to events. Conversely, dosage errors at a drug bottling plant are far less likely to yield highquality mentions. The crowd-generated data literature has noted that wise crowds must have sufficient information and knowledge to provide quality insights <ref type="bibr">(Surowiecki 2005</ref><ref type="bibr">, Sunstein 2006</ref>). Future research could examine how the amount of knowledge and available information impacts users' quantity and quality of contributions to VoC channels.</p><p>3. Precision of Crowd-Generated Data: The viability of user-generated content channels comes with caveats. Willingness to disclose information via social media channels, as well as access to and usage of such channels in general, could result in signal sampling biases <ref type="bibr">(Anderson and</ref><ref type="bibr">Agarwal 2011, Abrahams et al. 2015)</ref>. The social media channels examined in this study garnered higher recall but lower precision relative to search. Hence, they embody greater signal and noise as a result of the competing effects of greater salience on one hand and credibility implications on the other. Furthermore, because certain channels such as search and Twitter may have a degree of separation from event detection tasks in terms of their primary use cases, sole reliance on such channels would not be prudent, as noted by recent issues with the Google Flu monitoring system <ref type="bibr">(Agarwal and</ref><ref type="bibr">Dhar 2014, Lazer et al. 2014)</ref>. However, the crowd-generated data literature emphasizes the importance of having robust signal aggregation methods <ref type="bibr">(Surowiecki 2005)</ref>, and other studies have observed that these issues are preventable by including appropriate contextualization mechanisms <ref type="bibr">(Broniatowski et al. 2014</ref><ref type="bibr">, Brynjolfsson et al. 2016)</ref>.</p><p>Our study touched upon the potential for signal fusion. An important future direction is to explore more precise signal fusion methods <ref type="bibr">(Adjeroh et al. 2014)</ref>. Crowdsourcing methods that can enhance signal-tonoise ratios may also offer enhanced detection capabilities <ref type="bibr">(Brynjolfsson et al. 2016</ref>). Nevertheless, the results presented in this note constitute an important first step in understanding adverse event detection via crowd-generated data.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="8.2.">Contributions to Practice</head><p>Risk management groups and IT departments can use the analysis framework and empirical insights to develop their VoC listening platforms in a more rigorous, systematic manner <ref type="bibr">(Abbasi et al. 2018</ref><ref type="bibr">, Kitchens et al. 2018)</ref>. Our framework suggests that practitioners begin by understanding their key monitoring objectives, including which event types they wish to monitor. The ensuing channel and detection methods investment decisions can be driven by the nuanced precision, recall, and timeliness implications of their environment. For instance, we noted earlier how differences in the quantity of products being monitored and available monitoring resources could result in varying perspectives on precision, recall, timeliness trade-offs for regulators versus individual firms. Furthermore, the results of our GASD method suggest that robust event detection methods applied to appropriate channels have the potential to offer timely detection with manageable false-positive rates, making enterprise VoC listening feasible. Collectively, by addressing many of the key impediments to VoC listening platform adoption and business value <ref type="bibr">(Browne et al. 2015</ref><ref type="bibr">, Davies 2015)</ref>, this study has the potential to enhance outcomes related to practitioner's VoC listening platform investment decisions.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="8.3.">Limitations</head><p>Our work is not without its limitations. Timeliness is a relative construct: as the status quo changes and the state of the art advances, what is considered timely today may not be in the future. Although omnichannel VoC listening platforms are intended to alleviate some of the availability biases inherent in spontaneous reporting databases, they are not entirely devoid of such biases. Certain products might be better suited to online monitoring. Moreover, disparities such as literacy and socioeconomic factors could moderate the frequency and salience of crowdgenerated signals. Consequently, it is conceivable that precision, recall, and timeliness could vary across firms or classes of products, creating potential inequities. From an ethical standpoint, examining adverse event detection biases attributable to, or amplified by, the use of machine learning approaches applied to usergenerated content constitutes an important future research direction. Similarly, a deeper exposition into multichannel listening strategies that examine broader stakeholder scenarios and consider additional fusion methods and channels constitutes an important future direction.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>Information SystemsResearch, 2019, vol. 30, no. 3, pp. 1007-1028, &#169; 2019 The Author(s)   </p></note>
		</body>
		</text>
</TEI>
