<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>On Mitigating Acoustic Feedback in Hearing Aids with Frequency Warping by All-Pass Networks</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>09/15/2019</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10351868</idno>
					<idno type="doi">10.21437/INTERSPEECH.2019-3195</idno>
					<title level='j'>Proceedings Article published 15 Sep 2019 in Interspeech 2019</title>
<idno></idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Ching-Hua Lee</author><author>Kuan-Lin Chen</author><author>Fred Harris</author><author>Bhaskar D. Rao</author><author>Harinath Garudadri</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Acoustic feedback control continues to be a challenging prob- lem due to the emerging form factors in advanced hearing aids (HAs) and hearables. In this paper, we present a novel use of well-known all-pass filters in a network to perform frequency warping that we call “freping.” Freping helps in breaking the Nyquist stability criterion and improves adaptive feedback can- cellation (AFC). Based on informal subjective assessments, dis- tortions due to freping are fairly benign. While common ob- jective metrics like the perceptual evaluation of speech quality (PESQ) and the hearing-aid speech quality index (HASQI) may not adequately capture distortions due to freping and acoustic feedback artifacts from a perceptual perspective, they are still instructive in assessing the proposed method. We demonstrate quality improvements with freping for a basic AFC (PESQ: 2.56 to 3.52 and HASQI: 0.65 to 0.78) at a gain setting of 20; and an advanced AFC (PESQ: 2.75 to 3.17 and HASQI: 0.66 to 0.73) for a gain of 30. From our investigations, freping provides larger improvement for basic AFC, but still improves overall system performance for many AFC approaches.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>This research is part of the Open Speech Platform (OSP) <ref type="bibr">[1,</ref><ref type="bibr">2,</ref><ref type="bibr">3,</ref><ref type="bibr">4]</ref> funded by an NIH/NIDCD initiative to enable psychophysical research beyond what is currently capable in support of hearing healthcare. This paper is related to improving acoustic feedback reduction for form-factor accurate, audiologic research in the field using behind the ear, receiver in the canal (BTE-RIC) transducers, hardware, embedded software, and application software we developed <ref type="bibr">[3,</ref><ref type="bibr">4]</ref>. In order to compensate for mild to moderate hearing loss, commercial hearing aids (HAs) and OSP provide an average gain of 35-38 dB. In the emerging form factors for advanced HAs and hearables, including conventional BTE-RICs, there is a significant acoustic coupling between the microphones and loudspeakers (called receivers in the telephony and HA communities). This acoustic coupling varies significantly based on surroundings (e.g. hats, scarves, hands, and walls that come in close proximity to the transducers) and can cause the system to become unstable, when the audio content includes characteristic frequencies of the system. This instability results in brief "howling" artifacts and they are of immense annoyance to the HA users.</p><p>Howling artifacts manifest when multiple factors collude to fulfill the magnitude and phase conditions of the Nyquist stability criterion (NSC) <ref type="bibr">[5]</ref>. Adaptive feedback cancellation (AFC) has been the work horse for breaking NSC to avoid instabilities in many audio applications <ref type="bibr">[6]</ref>, including HAs <ref type="bibr">[7,</ref><ref type="bibr">8,</ref><ref type="bibr">9,</ref><ref type="bibr">10]</ref>. Typically, the AFC deploys the least mean square (LMS) based approaches to mitigate the magnitude condition in NSC <ref type="bibr">[11,</ref><ref type="bibr">12,</ref><ref type="bibr">13,</ref><ref type="bibr">14]</ref>. On the other hand, frequency shifting (FS) <ref type="bibr">[15,</ref><ref type="bibr">16,</ref><ref type="bibr">17]</ref> and other ad hoc methods <ref type="bibr">[18,</ref><ref type="bibr">19,</ref><ref type="bibr">20]</ref> mainly deal with the phase condition. In this paper, we focus on spectral manipulations following LMS based approaches to break NSC in both magnitude and phase conditions.</p><p>In a 1972 paper, Oppenheim and Johnson <ref type="bibr">[21]</ref> described discrete representation of continuous signals and systems and included detailed recipes to "transform the frequency axis in a nonlinear manner." This frequency warping is accomplished using an all-pass network. The authors presented three applications (efficient spectral analysis with unequal resolution, vernier spectral analysis, and correcting for helium speech artifacts in underwater diving) and predicted additional applications in future. Surprisingly, this clever trick of using all-pass structures for manipulating signals seems to have attracted much less attention than warranted in the speech community.</p><p>We adopt teachings in <ref type="bibr">[21]</ref> for HAs and call it "freping," a portmanteau for frequency warping. A common type of hearing loss is the sloping hearing loss, where the impaired user has limited ability to perceive high-frequency content. Typically, the intervention is to boost the high-frequency components or move the content to lower frequencies <ref type="bibr">[22]</ref>. The former introduces challenges for acoustic feedback control, while the latter facilitates better feedback reduction. Another less common type of hearing loss, but more challenging for providing meaningful interventions is the "cookie bite" hearing loss, wherein it is difficult for the impaired person to perceive mid-frequency content, compared with low-and high-frequency components. We posit that freping will provide an additional tool to the audiologist for managing individual hearing loss profiles. In this paper, we focus on the benefits of freping for mitigating NSC in conjunction with LMS based AFC approaches.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Revisiting all-pass networks</head><p>The all-pass networks described in <ref type="bibr">[21]</ref> realize a nonlinear mapping of the frequency axis as controlled by a single warping parameter &#8629;. Let ! = 2&#8673;(f/fs) be the normalized angular frequency where f is the original frequency and fs is the sampling rate. The mapping &#10003;(&#8226;) is according to <ref type="bibr">[21]</ref>:</p><p>where ! = 2&#8673;( f/fs) and f is the warped frequency.</p><p>It can be shown that the nonlinear frequency mapping (1) between the original signal v(n) and the frequency-warped signal q(k) can be achieved by passing the time-reversed signal v( n) through a linear time-invariant system H k (z) given as:</p><p>and taking the output of H k (z) at n = 0 as q(k). It can thus be implemented as the network shown in Figure <ref type="figure">1</ref>. The first two stages act as (i) low-pass filters when &#8629; is positive and the network warps frequencies higher and (ii) high-pass filters when &#8629; is negative and the network warps frequencies lower. The remaining stages realize the actual frequency warping based on the bilinear transformation <ref type="bibr">[23]</ref>. Note that when &#8629; = 0, it simply passes through the input without any spectral modifications. The frequency-warped output is given by sampling e q k (n), the output signal at the k-th stage, along the cascade chain at n = 0, i.e., q(k) = e q k (0). In other words, the input sequence is first flipped and then passed through the network; the last sample of the output sequence at the k-th stage is taken as the k-th sample of the final frequency-warped sequence <ref type="bibr">[24]</ref>.</p><p>The all-pass network for frequency warping.</p><p>It is worth noting that in practice we need to truncate the signal for the all-pass network to be realizable. Therefore, the warping performance will depend on other factors such as the length and the type of the window function used.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Freping: real-time frequency warping</head><p>The all-pass networks described above are adopted for real-time frequency manipulations as illustrated in Figure <ref type="figure">2</ref>. The input signal is first divided into overlapping frames and windowed using a proper window function. Each windowed segment then goes through the all-pass network to perform frequency warping with a specified warping parameter &#8629;. Finally, the overlap-add method <ref type="bibr">[25]</ref> is applied to produce the frequency-warped signal.  To allow a more flexible way of manipulating spectral characteristics, we propose the multichannel freping as illustrated in Figure <ref type="figure">3</ref>. The system utilizes a set of band-pass filters (BPFs) which divide the input signal into M frequency bands and a set of warping parameters &#8629; = [&#8629;1, ..., &#8629;M ] T . Each band goes through an independent all-pass network with the corresponding warping parameter. The output signals of all the frequency bands are summed up to produce the frequency-warped signal.</p><p>In many practical situations, it is convenient to reuse the multichannel compression modules <ref type="bibr">[26]</ref> in HA processing for freping. For specific types of hearing loss (e.g. sloping, cookiebite, etc.), increasing the gain in higher frequency bands aids to fulfill the magnitude condition of NSC and freping hinders the phase condition to occur. Thus, freping provides a way for simultaneously optimizing the parameters of multichannel compression and frequency lowering <ref type="bibr">[27]</ref> in HAs for individual hearing loss. In this work, we limit ourselves to negative values of &#8629; so that freping always shifts spectral content lower. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Freping for acoustic feedback reduction</head><p>We investigate benefits of freping for mitigating acoustic feedback along with LMS based AFC, with the motivation of improving feedback control in the systems described in <ref type="bibr">[1,</ref><ref type="bibr">2,</ref><ref type="bibr">3,</ref><ref type="bibr">4]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.">Adaptive feedback cancellation (AFC) system</head><p>We adopt the AFC framework used in <ref type="bibr">[14]</ref> as depicted in Figure <ref type="figure">4</ref>. The AFC filter W (z, n), placed in parallel with the HA processing G(z, n), is the transfer function of an L-tap adaptive filter w</p><p>T that continuously adjusts its coefficients to capture the time-varying nature of the acoustic feedback path F (z, n). d(n) is the microphone input which contains the clean signal x(n) and the feedback signal</p><p>is the feedback-compensated signal. A(z, n) is a time-varying pre-filter to decorrelate the input and output signals based on the prediction error method (PEM) <ref type="bibr">[7]</ref>. B(z) is a band-limited filter to concentrate on the frequency region where oscillation is more likely to occur <ref type="bibr">[11]</ref>. Typically, LMS-type algorithms are carried out for coefficient adaptation using the pre-filtered signals u f (n) and e f (n) to update the AFC filter w(n) as:</p><p>where</p><p>is the step size parameter, &gt; 0 is a small constant to prevent division by zero, and</p><p>is the power estimate with a forgetting factor 0 &lt; &#8674; &#63743; 1. The update rule (3) is actually the "modified" LMS using the sum method <ref type="bibr">[28]</ref> and has been widely used in AFC works <ref type="bibr">[7,</ref><ref type="bibr">11,</ref><ref type="bibr">12,</ref><ref type="bibr">14</ref>]. An advanced AFC algorithm, based on the LMS (3), is the sparsity promoting LMS (SLMS) proposed in <ref type="bibr">[14]</ref> which leverages the sparsity of the feedback path impulse response to achieve faster convergence for improvement. The SLMS update rule includes an additional sparsity promoting term S(n) as:</p><p>where S(n) = diag{s0(n), s1(n), ..., sL 1(n)} is an L-by-L diagonal matrix and the diagonal elements are updated according to si(n</p><p>where p 2 (0, 2] is the sparsity control parameter and c &gt; 0 is a small positive constant to avoid stagnation of the algorithm.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2.">Mitigating Nyquist stability criterion (NSC)</head><p>Without any feedback control mechanism, the frequency responses of the HA processing G(e j! , n) and the feedback path F (e j!</p><p>, n) form a closed-loop system which exhibits instability that leads to howling. The NSC <ref type="bibr">[5]</ref> states that the closedloop system becomes unstable whenever the following magnitude and phase conditions are both fulfilled <ref type="bibr">[8]</ref>:</p><p>When AFC in employed, it becomes:</p><p>where F (e j!</p><p>, n) = B(e j! )W (e j! , n) is the estimated feedback path frequency response. The AFC aims at minimizing</p><p>, n) F (e j! , n) to mitigate the magnitude condition. It is well-known that the LMS-type algorithms widely used in AFC suffer from biased estimation due to signal correlation <ref type="bibr">[29]</ref>. Consequently, the feedback path estimate can be erroneous if decorrelation is not carefully considered. Although the PEM-based pre-filter <ref type="bibr">[7]</ref> has provided certain amount of decorrelation, further improvement is achievable by inserting additional signal processing into the forward path of the HA <ref type="bibr">[17]</ref>, usually placed at ? as shown in Figure <ref type="figure">4</ref>. Existing methods include frequency shifting (FS) <ref type="bibr">[15,</ref><ref type="bibr">16,</ref><ref type="bibr">17]</ref>, phase modulation <ref type="bibr">[18]</ref>, time-varying all-pass filters to introduce phase shifts <ref type="bibr">[19]</ref>, linear predictive coding vocoder <ref type="bibr">[20]</ref>, to name a few. In general, quality degradation might be introduced by these decorrelation methods and thus there is the trade-off between the sound quality and the decorrelation ability for AFC improvement.</p><p>Freping is an extreme version of FS <ref type="bibr">[22]</ref> and it plays a similar role for decorrelation. It introduces nonlinear frequency shifts and the distortions appear to be perceptually benign based on informal subjective assessments. As instability is most likely to occur at the high-frequency region, it is reasonable to manipulate the high-frequency content while keeping the low-frequency region intact to avoid degradation in quality. By providing additional decorrelation, freping can reduce the AFC bias and thus a better feedback path estimate can be obtained, thereby improving the magnitude condition in NSC. On the other hand, freping also helps avoid the microphone and receiver signals from remaining continuously in phase with each other. This prevents the phase condition in NSC to hold at the same frequency at two consecutive instants. Consequently, the input and output sounds could not build up in amplitude as effectively. Therefore, the likelihood of instability is reduced.</p><p>Note that the approach in <ref type="bibr">[19]</ref> also utilizes all-pass filters to achieve decorrelation, in which time-varying poles are used for introducing phase shifts. This is different from freping which manipulates the spectral magnitude as well. Since freping is similar to the FS, we compare them in the following section.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Evaluation</head><p>We evaluate the proposed freping system using computer simulations in MATLAB at a sampling rate of 16 kHz. We implemented a 6-band system using a set of BPFs with non-uniform bandwidth whose center frequencies are 250, 500, 1000, 2000, 4000, and 6000 Hz, respectively. Frames of 128 samples with 50% overlap were utilized. The Hann function was applied for windowing. 25 male and 25 female speech signals from TIMIT database were used for simulations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1.">Speech quality considerations</head><p>In this experiment we directly performed freping on the speech signal and measured the frequency distortion at the output using the MATLAB implementation <ref type="bibr">[30]</ref> of the (wide-band) perceptual evaluation of speech quality (PESQ) <ref type="bibr">[31]</ref>. The PESQ score gives a good prediction of the mean opinion score and has been suggested for quantifying spectral distortion brought by FS <ref type="bibr">[29,</ref><ref type="bibr">8,</ref><ref type="bibr">18]</ref>. Figure <ref type="figure">5</ref> shows the average PESQ score of the freping output over the 50 speech files as a function of the warping parameter &#8629;, for the cases of operating on the full-band (&#8629; = &#8629;[1, 1, 1, 1, 1, 1] T ) and on the last two (high) frequency bands (&#8629; = &#8629;[0, 0, 0, 0, 1, 1] T ). We can see that quality degradation is minor in the latter case.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2.">Acoustic feedback reduction with freping</head><p>Now we consider the practical scenario of HA as in Figure <ref type="figure">4</ref>.</p><p>We study freping with &#8629; = &#8629;[0, 0, 0, 0, 1, 1] T on top of the LMS (3) and the SLMS (4). The experimental setup was as follows. The HA processing G(z, n) = gz 4 where g is the HA gain and 4 is the sample delay chosen to have a total HA latency under 10 msec (from d(n) to o(n)). The feedback path impulse response was measured using a BTE-RIC device with open fitting on a dummy head with a handset placed on the earthe most challenging scenario for breaking NSC. For the AFC, we used L = 100, &#181; = 0.005, &#8674; = 0.985, and = 10 6 for both LMS and SLMS. For the SLMS we used p = 1.5 and c = 10 6 as suggested in <ref type="bibr">[14]</ref>. In all simulations, the AFC filter coefficients were initialized as all zeros.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2.1.">HA output quality</head><p>Figure <ref type="figure">6</ref> shows the average PESQ score of the HA output over the 50 speech files for several values of the warping parameter &#8629;. From the results we can see that when we increase &#8629; in magnitude from 0, acoustic feedback gets better controlled, resulting in improved quality. However, further increasing &#8629; in magnitude leads to higher spectral distortion and thus the quality drops. This indicates the trade-off between the reduction of feedback artifacts and frequency distortion, and is better seen in the case of a more aggressive gain setting.  With &#8629; = 0.02, PESQ improvements of 2.56 to 3.52 and 2.75 to 3.17 can be seen for LMS with HA gain at 20 and SLMS with HA gain at 30, respectively.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2.2.">Feedback reduction improvement</head><p>We now focus on quantifying the improvement brought by freping in reducing feedback artifacts. In the remaining experiments, &#8629; = 0.02 was used as suggested by the results in Figure <ref type="figure">6</ref>. Also from Figure <ref type="figure">5</ref>, this choice of &#8629; corresponds to an average PESQ of 4.55 which indicates good quality. Furthermore, based on informal subjective assessments, distortions introduced with &#8629; in the vicinity of this choice are fairly benign.</p><p>According to <ref type="bibr">(1)</ref>, for this choice of &#8629;, the center frequencies of the fifth and sixth frequency bands would move from 4000 and 6000 Hz to 3898 and 5927Hz, respectively. We compare performance with an existing FS method based on the analytical representation of signal using the Hilbert transform <ref type="bibr">[6,</ref><ref type="bibr">15]</ref>. The amount of shift was set to 12 Hz, only applied to frequency region above 1.5 kHz as suggested by <ref type="bibr">[16,</ref><ref type="bibr">17]</ref>. When directly performed, this setup gives an average PESQ score of 4.47 of the FS output over the 50 speech files, which is comparable but slightly lower than that of the freping result.</p><p>For evaluation, we compare the feedback-compensated signal e(n) with the clean signal x(n), using the hearing-aid speech quality index (HASQI) <ref type="bibr">[32]</ref> which has been adopted in prior AFC works <ref type="bibr">[14,</ref><ref type="bibr">10,</ref><ref type="bibr">33]</ref>. The HASQI score ranges from 0 to 1, where a higher value indicates better quality.</p><p>Figure <ref type="figure">7</ref> presents example spectrograms of the feedbackcompensated signal for several cases. We can see that freping effectively reduces the howling components present in the red boxes, resulting in improved quality.</p><p>Figure <ref type="figure">8</ref> demonstrates advantage of using freping by showing the average HASQI score over the 50 speech files for various gain settings. From the results we see that both the basic (LMS) and advanced (SLMS) AFC algorithms can benefit from freping. This indicates the ability of the proposed frequency warping method to further improve feedback reduction on top of many AFC approaches. Moreover, compared to the FS, freping demonstrates better performance under all the gain settings.</p><p>Finally, we compare the added stable gain (ASG), which is the additional gain due to feedback control mechanism that the HA can still operate in the stable state, for the cases of AFC,  AFC with FS, and AFC with freping. We used the ASG estimation approach proposed in <ref type="bibr">[10]</ref>, where a HASQI below 0.8 was considered of unacceptable quality. The results are shown in Table <ref type="table">1</ref>, obtained from the average of 5 male and 5 female speech files. We can see that freping can improve the ASG on top of both the basic and advanced AFC algorithms. Compared to the FS, a higher ASG can be achieved by using freping. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.">Conclusions</head><p>In this paper, we proposed a novel use of all-pass networks for frequency warping that we call "freping." We described realtime realization of multichannel freping for use in HAs and its use for breaking the NSC in acoustic feedback control. Experimental results demonstrate quality improvements with freping for basic and advanced AFC approaches. For a desired quality lower bound (e.g. HASQI = 0.8), we found ASG improvements of 2.5 and 1.4 dB for LMS and SLMS with freping, respectively.</p></div></body>
		</text>
</TEI>
