How CarmaDeal calculates Community Wisdom
Methodology 3.0. Owner sentiment is calculated only from unique, vehicle-matched current or former owners. Structured ratings, recommendation responses, and would-buy-again answers receive more weight than inferred text sentiment. Posts from the same owner are combined so each owner receives one vote.
NHTSA complaints, recalls, investigations, technical bulletins, expert reviews, complaint databases, and AI-generated estimates do not affect owner sentiment. They are analyzed separately to identify reliability, defect, safety, and inspection concerns.
Records are matched by year, generation, trim, engine, transmission, body style, and market. Records involving materially different versions of the vehicle are excluded. Evidence is grouped by generation: model years that share a confirmed platform (every 2017–2021 Civic Type R is the FK8) count as the same car, with exact-year and same-generation counts always disclosed separately.
Low confidence does not pull a score toward average. CarmaDeal either shows the estimated score with an uncertainty range or withholds the score when the qualified owner sample is insufficient.
1 · Source classification & eligibility
Every record gets a source type before it can enter any calculation. Eligibility is a hard table — the model cannot override it:
| Source type | Owner sentiment | Owner reliability | Issue detection | Official safety |
|---|---|---|---|---|
| structured owner review | owners only | owners only | yes | no |
| carmadeal owner survey | owners only | owners only | yes | no |
| owner forum post | owners only | owners only | yes | no |
| owner reddit post | owners only | owners only | yes | no |
| owner social post | owners only | owners only | yes | no |
| owner video review | owners only | owners only | yes | no |
| journalist review | never | no | no | no |
| expert blog | never | no | no | no |
| nhtsa complaint | never | no | yes | yes |
| nhtsa recall | never | no | validation only | yes |
| nhtsa investigation | never | no | validation only | yes |
| manufacturer tsb | never | no | validation only | no |
| complaint aggregator | never | no | yes | no |
| repair database | never | no | no | no |
| unknown | never | no | no | no |
Hard rule: complaint channels are self-selected and negative by construction — nobody files an NHTSA complaint to praise their car. They inform issue and safety sections only.
2 · Owner status & one-owner-one-vote
Every potential sentiment record is classified by ownership status, with the verbatim evidence span retained (an ownership claim without verifiable evidence carries no weight — and "verified owner" can only come from platform verification, never from the model):
| Owner status | Weight |
|---|---|
| verified owner | 1.00 |
| explicit owner | 0.80 |
| former owner | 0.80 |
| likely owner | 0.50 |
| prospective buyer | 0.00 |
| test driver | 0.00 |
| non owner | 0.00 |
| unknown | 0.00 |
Records cluster into unique owners by platform author identity and duplicate-text detection. Cross-posts, quoted content, and multiple updates from the same owner merge into one owner-level record — the owner, not the post, is the unit of analysis.
3 · Vehicle matching
The requested vehicle is normalized into a canonical object (year, make, model, generation, trim, engine, transmission, body style, market — via an alias dictionary of chassis codes and variant facts). Every record is then matched:
| Match level | Base weight |
|---|---|
| exact | 1.00 |
| generation | 0.60 |
| model family | 0.00 |
| make only | 0.00 |
| mismatch | 0.00 |
| unknown | 0.00 |
Evidence is grouped by generation. Model years that share a platform are the same car — every 2017–2021 Civic Type R is the FK8 — so when the generation span comes from the dictionary, a same-generation owner report counts at full weight (1.00), not the reduced base weight. The 0.60 discount applies only when the generation window is an approximation (no dictionary entry) or the record never establishes which variant it describes. Exact-year and same-generation owners are always reported as separate counts.
Mismatched variants (wrong engine, wrong transmission, wrong body style, a different trim) are excluded outright and listed in the report's evidence drawer with the reason.
4 · The owner score
Each unique owner gets one score from the signals they actually expressed — nothing is invented for them:
| Signal | Weight (renormalized over available signals) |
|---|---|
| normalizedRating | 0.45 |
| wouldBuyAgain | 0.30 |
| wouldRecommend | 0.15 |
| textSentiment | 0.10 |
Text-only records carry a reduced quality multiplier (0.7). Each owner's evidence weight is ownerStatus × vehicleMatch × ownershipDepth × sourceQuality; first-day impressions weigh 0.30 while 12-month/10k-mile ownership reports weigh 1.00. Calendar recency is not a universal penalty — a detailed five-year review does not expire.
5 · Source caps, the final score, and honest uncertainty
With exactly two source families, each may contribute at most 50% of total weight. Once three or more families are present, no family may contribute more than 40%; excess redistributes to the other families. Confidence comes from a weighted bootstrap over owner clusters (2,000 iterations, deterministic seed, caps reapplied per iteration) reported as a 95% interval.
No shrinkage, ever. The score is never pulled toward 50. It is withheld (with the reason shown) when any of these hold: fewer than 10 unique qualified owners · fewer than 2 owner-source families · majority likely-owner weight · sub-threshold vehicle-match confidence · primarily text-only signals · excessive source concentration.
Confidence tiers: Insufficient (<10 owners — no numeric score) · Low (10–24, 2+ families) · Medium (25–74, 3+ families, 70%+ vehicle-confirmed weight, 50%+ explicit/verified owners) · High (75+, 4+ families, 85%+ vehicle-confirmed, 70%+ explicit/verified, no source over the cap). "Vehicle-confirmed" means the exact year or a dictionary-confirmed same-generation match — grouped by generation, they are the same car.
6 · Owner-reported reliability (separate from sentiment)
A person can love a car that broke, or dislike a car that never did. Reliability uses only comprehensive ownership records (stated duration/mileage or concrete issues) — a short "love it" post is never counted as proof of zero problems:
| Issue severity | Value (0–4) |
|---|---|
| cosmetic | 1 |
| minor functional | 1 |
| moderate repair | 2 |
| major drivability | 3 |
| loss of control | 4 |
| crash fire risk | 4 |
The result is the reported risk in the analyzed owner sample — never a claim about the failure rate of all vehicles.
7 · Issue signals, timeline, and official data
Issue frequency is reported as evidence signals, not fabricated percentages — self-selected samples have no valid population denominator:
| Signal level | Unique qualified owners |
|---|---|
| isolated | 1+ unique qualified owners |
| emerging | 2+ unique qualified owners |
| recurring | 5+ unique qualified owners |
| well established | 10+ unique qualified owners |
Official NHTSA corroboration (complaints, recalls, investigations) is tracked separately from owner counts. The mileage timeline uses only owner-stated mileage (3+ owners per issue, 5+ total) — manufacturer maintenance intervals and industry repair estimates are labeled separately, never mixed in.
Repair costs, maintenance schedules, depreciation, and ownership-fit narratives are modeled estimates, clearly labeled, kept out of every owner score.
8 · The optional composite
Published only when both components reach medium confidence, always with the components shown beneath it. The 60/40 split is a product-policy weighting, not a naturally occurring statistical constant. Safety ratings, recall counts, complaint volume, and AI estimates are never part of it.