One Side Rates, the Other Can't Answer
written by Stefan Christoph
- 21 minutes readA note on where this comes from: I’m not a sociologist or a platform designer. I read the research because I keep landing on both sides of the star โ some weeks I’m the one tapping out a rating without thinking, some weeks I’m the one refreshing a score that decides something real. Every study I lean on is cited at the bottom. Where a claim is my synthesis rather than a finding, I say so. And nothing here is legal advice; defamation, review removal, and the like show up only as categories, never as counsel.
Four marks, one direction
Picture four moments.
It’s the end of a ride. The driver was fine โ got you there, didn’t chat, took the route you’d have taken. You give three stars, or maybe you don’t rate at all, and you’re already thinking about the meeting you’re walking into. What you don’t see is the other side of that tap. On the platform in the study I’ll lean on most, a driver’s score is the average of their last 500 rated trips, and they need to hold around 4.6 out of 5 to keep working [1]. Your “it was fine” three didn’t feel like a verdict to you. In an average that has to stay above 4.6, a three is a failing grade. You experienced an opinion. He experienced management.
It’s a Tuesday for a small restaurant. The owner opens the laptop and there’s a new one-star: slow service, cold food, a name she doesn’t recognize on a night she remembers being slammed and short-staffed. She knows what happened. A walk-in of twelve with no notice, a line cook who’d called in sick. She cannot say any of that. The review keeps the accusation; it has no field for the circumstances. She can reply, carefully, once, in a tone the platform and her future customers will judge, but she cannot cross-examine, and she cannot make it go away. The mark stays, searchable, folded into the number that decides who walks in next month.
It’s a reference call the candidate will never hear. Somewhere, a former manager is being asked the quiet questions โ would you hire her again, how did she handle pressure โ and the answers, warm or cool or carefully hedged, will move an outcome the candidate has staked a year on. She will never know it happened. There is no reply channel at all. The most consequential rating in a career is often the one delivered entirely offstage.
And it’s a family, over years. Somewhere in the shared story there’s a line about one person โ “she’s the difficult one,” “he’s unreliable” โ assembled from a handful of moments, most of them old, none of them cross-examined, and quietly updated at every gathering. It gates real things: who gets asked, who gets trusted, who gets the benefit of the doubt. The person it’s about can feel it in the room and can’t quite answer it, because there’s no moment where the verdict is read aloud and a defense is invited. It’s just the weather now. (That’s an example, not a claim about anyone’s family in particular โ but most of us can find the version that fits ours.)
Four marks, four settings, and only one of them involved a business. In each, one side made a lasting, aggregated judgment casually, and the other side lives inside it with almost no way to answer. The star, the review, the reference, the family verdict โ same shape. That shape is the eighth entry in this series.
The asymmetry, again
I’ve been collecting tables where two people live in different worlds. Part 7 was about reversibility: the door that locks behind only one of you, the mark that’s a one-way for the person it’s on and a shrug for the person who made it [2]. This post is that permanent mark seen from the other side. It can’t be undone, and, worse, it was one-directional to begin with. One side speaks into the record; the other can barely speak back.
Here’s the shape. Classic reputation was mutual and local: in a village, both sides could talk about each other, and everyone knew the source. Modern reputation is one-directional, global, permanent, and aggregated. The passenger rates the driver; the guest rates the restaurant; the former employer “rates” through a reference, and the rated party’s reply channel is weak, monitored, or missing entirely. Three things happen at once, and each makes the gap wider:
- One direction. The rater marks; the rated mostly cannot mark back, or cannot in any way that counts. The waiter does not get to rate the guest in a way that follows the guest around.
- Permanence. The mark outlives the moment. It’s searchable, it doesn’t decay on its own, and it’s the one-way door from Part 7.
- Aggregation. The step that turns anecdote into a gate: a pile of casual, unaudited, mood-carrying inputs is averaged into a single number that decides future access โ visibility, employment, the algorithm’s favor.
The asymmetry is invisible until the rated side can feel it: feel that a stranger’s idle tap is now part of the number that runs their livelihood, and that there’s no counter to file. And here’s the reframe that carries the whole post, the same one that ran under the earlier parts: most of the time, nobody is being cruel. The passenger tapping three stars is not savoring anyone’s deactivation. They are power-illiterate, not malicious; they genuinely do not know what the tap does. So let me take the two sides one at a time, because they have completely different problems.
Two sides of the same star
The person doing the rating
You experience your rating as an opinion. That’s the trap, because it isn’t only an opinion โ it’s an input to someone’s management, and you almost never see that half.
Three things are hidden from you. Your power: a casual three-star materially damages someone whose score has to stay high; you feel a shrug, they feel a strike. Your bias: whatever mood, hurry, or prejudice you brought to the moment rides straight into their record, with no calibration and no appeal โ more on that when we get to the research, because it’s the ugliest part. The context you don’t have: the rated party usually can’t explain the thing you’re marking them down for. The driver couldn’t tell you the previous passenger made him late; the restaurant can’t tell you about the walk-in of twelve. The record keeps your accusation and drops their circumstances.
Your failure mode is rating on autopilot โ marking the person for the system’s fault, punishing a bad night as if it were a bad character, and feeding a number you’ve never been taught to read.
The person being rated
You experience the rating as management, because it is. Your problem is that it’s permanent, aggregated, and mostly unanswerable, and the rational responses to that all have costs.
You can reply, but not freely: professional norms, platform rules, and plain power constrain what you can say, and the person you’re really “answering” (the angry author) has usually stopped listening. You can try to educate raters, plead for stars, over-serve defensively, the behavior the research calls playing “the rating game,” where surveillance becomes self-discipline [3]. And you carry the mark with the volume turned up: a public negative lands several times harder than a positive (the negativity bias we’ll get to), you reread it while the rater forgot writing it, and answering back honestly can cost you the very score you’re trying to protect. Whoever needs the number more has to swallow more, the least-interest principle, applied to reputation.
The thing connecting the two sides is one question of translation: can the rating side learn what the tap actually does, and can the rated side answer the record instead of the reviewer? To see why both are hard, we go to the research. This is the long middle of the post, and it’s the why behind every move in the playbook; if you only want the what, jump ahead โ the playbook will be waiting.
What the research actually says
The usual caveat, because it sets how much weight this bears. The Uber work is a qualitative case study; the reputation and negativity findings are a mix of experiments, surveys, and field observation. It’s real signal (the mechanisms are solid enough to plan around) but the exact magnitudes are softer than clean prose makes them sound. And to be clear about the framing: the platform study below is evidence, not a target. The point isn’t that one company is villainous; it’s that a design almost everyone now uses has predictable effects, and the effects show up in references and reviews and family stories just as much.
Complaint was supposed to trigger repair โ a review doesn’t
Albert Hirschman gave us the frame in 1970. When you’re dissatisfied, you have two basic moves: exit (leave) and voice (speak up and try to change things), and which you pick is moderated by loyalty [4]. The crucial part for us is what voice was for: it fired inside an ongoing relationship, so the complaint could trigger repair. You tell the shopkeeper the bread was stale; the shopkeeper fixes it; the relationship continues, improved.
Ratings broke that loop. A review is a third thing Hirschman didn’t quite have: exit with a permanent mark. It’s voice that fires on the way out, after the relationship is over, with no repair phase: all verdict, no dialogue. The feedback that was supposed to help someone improve instead becomes a scar they can’t respond to. That single structural change, detaching the complaint from the repair, is most of what makes modern reputation feel unfair. It kept the sting of voice and threw away its point.
Consumer ratings are quiet management
Alex Rosenblat and Luke Stark studied Uber drivers and found the thing that reframes the whole rating economy: the passenger rating system turns customers into de-facto managers of workers [1]. Their term for the broader machinery is algorithmic management โ the app implements the policies a human boss used to. But the ratings piece is sharper than that. The passenger, they write, becomes a stand-in for a traditional manager, “the watcher,” with the power to lower someone’s score and, through it, threaten their access to the job. Drivers are averaged over their last 500 rated trips and need to hold around 4.6 to stay active; some get deactivation notices if the last 25 or 50 trips dip [1].

Management, delivered as a tap. The passenger experienced an opinion at a red light; the driver reads the number that decides whether he keeps working next month โ averaged, permanent, and unanswerable.
Now sit with the asymmetry the study names directly: information. The person being managed by the number often can’t see how it’s computed, which trips count, or why it moved. A driver in the study puts it exactly: “What’s the formula to get these results?” [1]. They report their score dropping even when their behavior didn’t change, and being unable to get an unfair rating removed. This is management without any of the things that make management legitimate โ no calibration, no explanation, no appeal, no HR. The customer would never be allowed to do this through a formal channel. The rating pipe delivers it unaudited.
The pipe launders bias
Here’s the ugliest consequence, and it follows directly. If customer ratings are management, then customer bias is now an input to employment. Rosenblat and colleagues, in “Discriminating Tastes,” argue that consumer ratings become vehicles for exactly the demographic and personal biases that anti-discrimination rules exist to keep out of hiring and firing [5]. A manager who marked people down for their accent or their name would be a legal problem. A crowd of customers doing the same thing, one casual star at a time, produces the same outcome with no fingerprints โ the bias is distributed, deniable, and baked into an average. No single rater sees themselves as the discriminator. The aggregate does the discriminating.
I want to be careful here, because this is a correlational and structural argument, not a claim that every low rating is bias. The point is narrower and still serious: the rating channel has no filter for bias, so whatever bias exists in the crowd flows straight through to someone’s livelihood. That’s a design fact, not an accusation about any particular rater.
The medium runs on negativity
Two well-established tendencies make the aggregate worse than the underlying experience. First, negativity bias: across a wide body of work, bad is stronger than good โ negative events, emotions, and information weigh more heavily and stick longer than positive ones of the same size [6]. A one-star lands harder on the reader (and on the rated party) than a five-star lifts.
Second, and this is where I’m partly synthesizing rather than citing a single number, who bothers to write. The motivated tend to be the aggrieved: outrage writes the review, satisfaction scrolls on. Put those together and the distribution a reader sees skews more negative than the actual experience base โ not because the product is bad, but because the medium samples for anger and then weights anger. I’ll flag that as my reading of the negativity and reviewing literature rather than a measured constant; the direction is well-supported, the exact size varies by platform and study. Either way, the rated party is defending against a sample that’s stacked twice against them.
How one number gets made. A broad base of real experiences is filtered by who bothers to write, weighted so the negative dominates, then compressed into a single score that gates access. Three distortions, one number, no cross-examination.
The rated party lives with a one-way door
Pull the strands together and you get the lived position of the rated side. The mark is quasi-irreversible (Part 7’s one-way door), searchable, and aggregated into a number that gates opportunity. It’s received with negativity’s amplifier turned up โ reread, replayed, weighed several times heavier than any single good rating. And the honest reply is suppressed by dependence: answer back too sharply and you risk the algorithm, the platform relationship, or your professional standing. The structure doesn’t just deliver the blow; it confiscates the defense.
The playbook
Enough mechanism. Here’s what to actually do on each side โ and then a short list for anyone who builds these systems, because that’s where the deepest fixes live. As before, notice how cleanly the moves transfer from the app to the reference call to the kitchen table.
If you’re the one rating
The rating side’s five moves. None of them is ‘rate nicer.’ They’re about knowing what the tap does and aiming it at the right target.
1 ยท Learn the scale’s real grammar
On most platforms, anything under five is a strike, not a nuance โ a “four” can be a failing grade in a system that deactivates below 4.6 [1]. Your five-point scale is not the school grading you’re picturing; it’s closer to pass/fail with decorations. Either rate inside the system’s actual grammar or abstain. A thoughtful three, in a school-grades frame, is often a wound you didn’t mean to inflict.
2 ยท Separate the person from the system
The most common injustice of the rating pipe is the misdirected mark: the courier punished for the warehouse’s delay, the driver for the app’s bad route, the support agent for the policy they didn’t write. Before you mark the person, ask whether the thing that annoyed you was actually theirs. Rate the system’s failures to the system, through whatever real feedback channel exists, and spare the human whose score you’re about to move.
3 ยท Write the review you’d say to their face
The distance the screen gives you is artificial; the mark is real. If you wouldn’t say it, in those words, to the person standing in front of you, don’t type it into a permanent record they’ll reread and you’ll forget. This isn’t about softening honest criticism โ it’s about closing the gap between the casualness of the tap and the weight of the consequence.
4 ยท Use voice before verdict
This is Hirschman, made personal. Before the permanent mark, give the fixable thing a chance to be fixed โ say something in the moment, ask for the manager, flag it while repair is still possible. Most operators would far rather fix a problem than be scarred by it, and most would never hear about it if your only move is the silent one-star on the way out. Voice inside the relationship is the repair loop the platform removed; you can put it back by hand.
5 ยท Rate the good, not only the bad
If you only write when you’re angry, you are personally feeding the skew โ you’re one of the motivated-aggrieved that tilts the whole distribution. The single counterweight the medium has is the occasional, deliberate positive from someone who’d normally just scroll on. Documenting a good experience isn’t fluff; it’s the only thing that keeps the aggregate honest.
If you’re the one being rated
The rated side’s five moves. Most are about a statistical, dignified defense โ making the average honest and refusing to be captive to one score.
1 ยท Answer the record, not the reviewer
Your public reply is not read by the angry author, who has moved on. It’s read by every future person deciding whether to trust you. So write for them: calm, factual, once. A dignified, specific reply to an unfair review is one of the strongest trust signals you have, not because it changes the reviewer’s mind, but because it shows the next hundred readers how you handle a hard moment. Never match the heat; the heat is the trap.
2 ยท Build volume โ it’s your statistical defense
The math that hurts you can also protect you. A single bad mark dominates a small sample and drowns in a large one. So make satisfied outcomes rate-able โ ask at the right moment, make it easy, catch people when they’re happy โ without ever paying for or faking a rating. You’re not gaming the number; you’re making it representative by giving the quiet-satisfied majority a reason to show up in the average that the angry minority already dominates.
3 ยท Don’t play the rating game past your dignity
The research is honest that the game is partly unwinnable by design โ servility and star-begging trade your long-term standing for marginal, temporary points [3]. So cap what you’ll pay. Decide in advance how much groveling is beneath the value of the work, and hold that line. Some scores aren’t worth the person you’d have to become to chase them.
4 ยท Keep your own record
Where appeals or disputes exist, the difference between an unanswerable mark and an answerable one is contemporaneous evidence. Note the disputed interaction โ date, what actually happened, who was there โ while it’s fresh. You can’t cross-examine a review, but you can sometimes turn “your word against theirs” into “your word against the record,” and the record is on your side if you kept one.
5 ยท Diversify your reputational surface
A single-score dependency is a single point of failure. If one platform’s number is the only thing standing between you and no work, you’ve handed a stranger’s tap far too much power. Own channels, direct relationships, portable proof of your work โ this is Part 5’s “keep alternatives,” pointed at reputation. The person who isn’t captive to one score is the person who can afford to answer it with dignity.
If you design or run these systems
The deepest fixes aren’t available to either side of the star โ they belong to whoever builds the pipe. The known mitigations are not mysterious; each just trades a little platform simplicity for fairness:
- Two-way visibility โ both parties rate, both see, so the power isn’t purely one-directional.
- Context fields โ let the rated party attach circumstances, so the record keeps more than the accusation.
- Response rights โ a real, protected reply channel that doesn’t risk the algorithm.
- Bias auditing โ check aggregates for the demographic skew “Discriminating Tastes” warns about, because the crowd won’t check itself [5].
- Decay โ let old marks fade, so a permanent record doesn’t gate a person forever on a bad month.
And a reminder that this isn’t only a consumer-platform problem: references and 360-degree reviews inside companies are rating systems too โ unaudited, aggregated, often unanswerable. The same design duties apply to the HR tool as to the app.
The same asymmetry, off the clock
Strip the platform away and the oldest one-sided rating system is just gossip. The absent party can’t answer; the negative travels better than the positive [6]; and the “aggregate” (someone’s standing in a group) quietly gates real opportunities, exactly like a score. When you’re the listener, you’re a rater who’s forgotten it: the story you’re hearing has no cross-examination, and your nod is a tap on someone’s record. The listener-side duty is real. Use voice before verdict with the person themselves; separate the person from the circumstance; and remember that a one-sided story is a one-star with better grammar.

The oldest rating system needs no app. The one person who could answer the story being told is the one who isn’t in the room.
Family narratives are the long-run version โ “she’s the difficult one,” assembled over decades from a handful of moments, updated at every dinner, gating who gets trusted with what. If you’re the one carrying an old family verdict, the rated-side moves transfer intact: answer the record calmly rather than fighting the reviewer, build a large volume of counter-evidence over time rather than one dramatic rebuttal, and don’t grovel for a rating from people who’ve stopped updating. And if you’re the one holding the verdict on someone โ a relative, a colleague, an old friend โ notice that you’re running an unaudited rating system with no appeals process, and that the humane thing is to let the mark decay.
The one line to keep
If you remember nothing else, keep the shape, because it holds at every star:
Ratings turned casual opinion into permanent, one-directional management โ so the fix is power literacy on one side and a dignified, statistical defense on the other. If you’re rating: learn what the tap actually does, separate the person from the system, and use voice before verdict. If you’re rated: answer the record not the reviewer, build the volume that makes your average honest, and don’t play the game past your dignity. And whoever builds the system owns the deepest fix โ give the rated side a way to answer back.
That’s the whole thing. Everything above is the evidence for why it works, and the naming that lets you reach for the right move when you’re the one tapping the star โ or the one living inside the average.
This is the eighth entry in the asymmetry series โ moments where two people share a table and live in different worlds. Part 7 was reversibility, the one-way door; this was reputation, the mark the door closes on. The lens keeps finding new rooms, and I’m not committing to a syllabus โ but there are more of these than I expected when I started.
For now I’d rather hear from you, from whichever side of the star you’re usually on. If you rate for a living, or just rate a lot: what’s your rule for when a low rating is fair versus when you’re really marking down the system behind the person? And if you carry a score that decides something real: what’s the reply, the ritual, the reframe that lets you answer the record without letting it own you? Put it in the comments. Someone reading this is about to tap three stars without thinking, and someone else is refreshing an average that decides their month โ and both would take the help.
Sources
[2] The Door Only Locks Behind One of You โ Asymmetry of Reversibility (Part 7)
[3] Alex Rosenblat โ The Rating Game: The Discipline of Uber’s User-Generated Ratings
[4] Albert O. Hirschman โ Exit, Voice, and Loyalty (overview)
[5] Rosenblat, Levy, Barocas & Hwang โ Discriminating Tastes: Customer Ratings as Vehicles for Bias
[6] Baumeister, Bratslavsky, Finkenauer & Vohs โ Bad Is Stronger than Good (2001)
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
๐ฌ Also available as a blog walkthrough video on YouTube
โค๏ธ Created with the support of AI (Kiro)