Listening to the Fed: Vocal Signals Across Chairs and Risk-Related Communication
Ana Lorena Jiménez-Preciado, Francisco Venegas-Martínez, Cesar Gurrola-Ríos, Ricardo Jacob Mendoza-RiveraHow does the Federal Reserve sound when it speaks? We measure the delivery of the Chair across 84 FOMC press conferences from 2011 to 2026, using Google’s Gemini 2.5 Flash to annotate 11,156 sixty-second segments spanning the Bernanke, Yellen, and Powell eras. The model returns ten paralinguistic variables, among them emotion, firmness, speech rate, and hesitation counts. Rather than trust these labels, we test them. Checked against regex-based transcript counts of disfluencies and against Praat, hesitations and speech rate prove reliable (Pearson r up to 0.92); firmness and emotion do not, showing little acoustic grounding and shifting when the prompt changes, so they are better read as sentiment drawn from the words than from the voice. Using the measures that survive this test, we find that calm delivery dominates, as one would expect of a Fed Chair, yet the three Chairs differ in ways that hold up under a hierarchical model, with the sharpest gaps in the unscripted question-and-answer session. Measured disfluency rises around episodes such as the 2022 Ukraine war, but not across the pandemic as a whole. The wider lesson is that multimodal models must be validated feature by feature before they are trusted as instruments.