Notes

Expression is preference plus the choice to speak

October 2, 2025

The comment I heard most from my dissertation committee was that social media measures expression while surveys measure preference. It is a clean line, and I liked it too. But a clean line is where the thinking starts, not where it ends, and I have spent a while now trying to say what it actually means.

My first instinct was to collapse the distinction. Surveys and social media both want the same thing, what a person thinks, so in some sense both measure preference. Social media just adds steps: the person has to run into the topic, decide it is worth saying something, and post. On that reading, expression is preference viewed through a noisier channel.

I no longer think that is right, and the correction is the useful part. Preference is what a question elicits. You ask, the respondent answers, and the instrument delivers the preference more or less directly. Social media adds something the survey removes by design: the person must choose to speak. So expression is not a noisier measurement of preference. It is preference plus a decision to express it.

Writing it that way makes the decision the object of study rather than an inconvenience. And it exposes the part a survey methodologist should worry about most: the decision to speak is not independent of the preference it would reveal. People with sharper views, or views they feel are under-represented, are more likely to post. The selection into expression is correlated with the thing being measured. That correlation is not noise around the signal. For anyone trying to read public opinion off social media, it is the measurement problem.

Once it is framed as a selection problem, two questions follow, and both are ones I want to keep working on.

The first treats expression as worth understanding in its own right. Why does anyone choose to post? One structure I lean on in the dissertation is that a shared event acts as a common stimulus: something happens, and it prompts a broad slice of the population to say something at roughly the same time. Treating the event as common is a real assumption, not a given, since exposure to an event is itself uneven. But to the extent it holds, it lets differences in who posts be read as differences in the decision to speak rather than differences in what people saw. That is the cleaner version of the mechanism I want to model.

The second treats expression as the route back to preference. If I can model why people choose to speak, I can use that model to correct for the selection, which is another way of saying I can specify a better calibration model. The point of the calibration is not the abstract adjustment for who happens to be on a platform. It is the concrete adjustment for who chose to speak about this.

So my current position is roughly where I started, but load-bearing rather than decorative. The expression-versus-preference distinction matters. It earns its place once you stop treating expression as degraded preference and start treating it as preference conditioned on a choice, then ask what governs the choice. That reframing is the first thread from the dissertation I wanted to put out in the open, and it is why this note exists.