E-commerce Product Answerability Score: A Self-Assessment for AI Recommendations
TL;DR
- Product answerability is not the same as AI visibility. It asks whether a product can be identified, compared, qualified against constraints, and represented with current evidence—not whether an answer system will retrieve or recommend it.
- Use critical fail gates before a weighted score. Variant ambiguity, unavailable public content, materially wrong price or inventory, unsafe eligibility claims, or fabricated evidence should stop the audit even when other sections look strong.
- Score nine dimensions on a transparent 100-point rubric. Identity, public availability, attribute completeness, fit, commercial freshness, evidence, comparison readiness, data consistency, and governance each answer a different operational question.
- Keep page, feed, schema, marketplace, review, and observed-answer facts separate. Agreement increases confidence; repetition across owned surfaces does not create independent proof.
- Category-specific attributes matter more than generic copy length. A laptop, skincare product, replacement part, food item, and sofa require different constraints, safety facts, compatibility fields, units, and exclusions.
- Do not average away failures. A catalog score of 82 can hide a top-selling variant with the wrong price or a regulated product with missing warnings. Roll up the distribution, critical fails, and revenue-weighted exposure separately.
- Turn the score into an action queue. Create content only when content is the missing layer; otherwise fix product data, merchant feeds, variants, rendering, evidence, policy, or the product itself.
The Decision This Score Should Help You Make
A VP Ecommerce does not need another “AI readiness” badge. The useful decision is whether a product is answerable enough to enter a governed observation panel, which missing facts should be repaired first, and which failure belongs to content, commerce operations, product data, technical delivery, evidence, or product fit.
The score should make the catalog easier to govern. It should not give a false numerical prediction about retrieval, ranking, citation, recommendation, conversion, or revenue.
Begin with one product decision
Choose a question such as “Can a buyer determine whether this hiking jacket fits wet-weather use under $200?” or “Can a repair technician verify that this part fits the declared model and year?” The audit is meaningless without the buyer, product, market, and constraint.
Decide whether the unit is a product or variant
Color may be cosmetic for one category. Size, material, voltage, memory, ingredient, pack count, seller, region, or condition may materially change the answer in another. Score the unit a buyer can actually purchase.
Route the failure to the right owner
A missing waterproof rating may be a product-data gap, an evidence gap, or a property the product does not have. Rewriting the description cannot legitimately create the fact.
When the fact exists but the team has placed it on the wrong kind of page, use the e-commerce page-type decision map to decide whether the PDP, category, buying guide, comparison, policy, support, or evidence page should own the answer.
| Executive question | Audit evidence | Possible decision |
|---|---|---|
| Can the item be identified? | Product, variant, seller, GTIN/MPN/SKU consistency | Fix entity/variant model |
| Can the item be accessed? | Public page, rendered facts, status, canonical | Fix technical delivery |
| Can it satisfy the prompt? | Attributes, fit, compatibility, exclusions | Fill facts or declare non-fit |
| Are commercial facts current? | Price, currency, stock, shipping, returns clocks | Fix commerce data pipeline |
| Is the claim supportable? | Documentation, tests, reviews, certificates | Add evidence or narrow claim |
| Can it be compared fairly? | Units, criteria, alternatives, unavailable states | Build comparison-ready units |
| Should it enter the panel? | Critical gates + score + confidence | Include, investigate, defer, exclude |
What Is Product Answerability?
Product answerability is the degree to which public, current, and attributable information can resolve a defined product question. It combines product identity, discoverable facts, constraint coverage, commercial truth, evidence, comparisons, data consistency, and governance.
Answerability is a source-side property
The audit examines whether the information environment contains a defensible answer. It does not inspect the private internals of an AI system or claim to know why one model selected one source.
Visibility is an observed outcome
Visibility requires a declared prompt, answer product, mode, market, time, eligibility rule, and coding method. A highly answerable product can be absent. A poorly answerable product can still appear with inaccurate or incomplete claims.
Recommendation adds fit
A product can be factually described yet be a poor fit for a given budget, body type, age, material preference, shipping window, safety requirement, device, or workflow. Correct exclusion can be a successful answer.
| Concept | Question | Evidence | Not established |
|---|---|---|---|
| Answerability | Can public facts resolve the decision? | Product/source audit | Retrieval or selection |
| Eligibility | Is the product valid for the prompt/panel? | Scope and gate rules | Visibility |
| Retrieval | Did a source enter the candidate set? | Observable citation/log where available | Recommendation |
| Mention | Did the product/brand appear? | Coded answer | Comparison or fit |
| Citation | Was a source attributed? | Coded answer/source | Positive treatment |
| Comparison | Was the item evaluated on criteria? | Coded answer | Preferred outcome |
| Recommendation | Was it selected for declared conditions? | Coded answer | Universal best status |
| Conversion | Did a measurable action occur? | Analytics/commerce record | Causality |
Fix the Audit Scope Before Scoring
A score can become flattering by quietly changing the unit. A page may describe a product family while the offer sells 24 variants across 3 sellers and 2 markets. The scope card stops that drift.
Declare the product entity
Record brand, product line, model, variant, bundle, condition, seller, identifier, market, language, and canonical URL. Include parent-child relationships and aliases.
Declare the buyer and job
“Best running shoes” is underspecified. A buyer recovering from an injury, shopping for trail use, requiring a wide size, or staying below a price threshold needs different evidence and exclusions.
Declare the observation clock
Use one timestamp and timezone for page, schema, feed, marketplace, checkout, and answer observations. Price and inventory can change between captures.
| Scope-card field | Synthetic example | Audit rule |
|---|---|---|
| Brand/product | Example TrailShell | Stable entity label |
| Variant | Women's M / blue / 2026 edition | Purchasable unit |
| Seller | Example Brand US | Seller-specific offer |
| Identifiers | SKU EX-TS-WM-B; GTIN example | Validate format/source |
| Market/language | United States / English | Align commercial facts |
| Buyer | Weekend hiker | Do not generalize to every user |
| Job | Waterproof shell under $200 | Declared decision |
| Required constraints | Size M, rain use, delivery in 7 days | Eligibility gates |
| Observation clock | 2026-08-02 16:00 PT | One aligned capture window |
| Audit version | PAS-1.0 | Reproducible rubric |
Run Critical Gates Before the 100-Point Score
Some defects make the total misleading. A product with polished copy and rich reviews should not pass if the selected variant has the wrong voltage, ingredient, eligibility, price, or stock state.
Identity gate
Fail when the purchasable entity cannot be distinguished from another model, variant, seller, or condition, or when material identifiers conflict without resolution.
Commercial-truth gate
Fail when displayed price, currency, stock, seller, checkout availability, required subscription, shipping eligibility, or return condition is materially wrong for the declared market and clock.
Safety and evidence gate
Fail when a material safety, age, allergy, dosage, compatibility, regulated, certification, warranty, or performance claim is unsupported, contradicted, or assigned to the wrong entity.
| Gate | Pass | Investigate | Critical fail |
|---|---|---|---|
| Entity | Product/variant/seller stable | Minor alias ambiguity | Wrong or conflated purchasable item |
| Public access | Current facts render publicly | Intermittent/partial | Unavailable, blocked, or fact hidden from target flow |
| Price/currency | Page, offer, checkout agree | Timestamp/region unclear | Material mismatch |
| Availability | Product can be bought as stated | Regional edge unresolved | In-stock claim cannot be purchased |
| Fit/safety | Required limitation visible | Evidence incomplete | Unsafe/false eligibility statement |
| Compatibility | Exact supported relationship | Version/model unclear | Wrong-device or wrong-part claim |
| Evidence | Source and entity match | Source quality uncertain | Fabricated/misattributed proof |
| Variant | Selected attributes attach correctly | Parent-child ambiguity | Facts borrowed from different variant |
Do not delete a critical fail to improve the average. Keep the fail, owner, source, clock, and remediation state in the record.
Use a Transparent 100-Point Rubric
The Product Answerability Score in this article is an illustrative decision rubric. It has not been validated as a predictor of AI ranking, citation, recommendation, sales, or revenue.
Score the evidence, not writing polish
A beautiful description cannot compensate for missing fit facts or a stale offer. Give points only when the item meets the documented requirement for the audit unit.
Add confidence separately
Score completeness and evidence confidence as different fields. A filled attribute sourced only from promotional copy has a different confidence state from a manufacturer specification, controlled test, or relevant independent source.
Apply gates after arithmetic
Calculate the raw score for diagnosis, then apply the gate status. A critical fail results in “not eligible for pass” rather than a cosmetically high final grade.
| Dimension | Points | Core question |
|---|---|---|
| Product/variant identity | 10 | Is the purchasable entity unambiguous? |
| Public availability/rendering | 10 | Can the important facts be accessed? |
| Attribute completeness | 15 | Are category-specific decision facts present? |
| Fit/constraint/exclusion coverage | 15 | Can the item be qualified responsibly? |
| Commercial facts/freshness | 15 | Are price, stock, shipping, returns, and offer current? |
| Evidence/provenance/reviews | 15 | Are material claims inspectable and attributable? |
| Comparison/alternative readiness | 10 | Can trade-offs be evaluated fairly? |
| Page/feed/schema consistency | 5 | Do machine-readable and visible facts agree? |
| Governance/change clocks | 5 | Can the truth stay current? |
| Total | 100 | Diagnostic total before gate status |
| Raw score | Illustrative interpretation | Gate override |
|---|---|---|
| 90–100 | Strong answer units; observe and maintain | Critical fail still blocks pass |
| 75–89 | Generally answerable with material gaps | Investigate all gates |
| 60–74 | Partial; prioritize high-value missing facts | Do not claim readiness |
| 40–59 | Fragmented or stale | Repair before broad panel use |
| 0–39 | Insufficient source truth | Rebuild data/evidence foundation |
These bands are workflow examples, not universal benchmarks.
Score Product and Variant Identity
AI answers, feeds, marketplaces, and human buyers can all misattribute a property when product names and parent-child relationships drift.
Normalize the entity graph
Record brand, product group, product, model, variant, seller, condition, identifiers, canonical URL, and allowed aliases. A model-year change may deserve a new entity even when the merchandising name stays similar.
Keep variant-determining properties visible
Size, color, material, pattern, memory, storage, voltage, flavor, quantity, condition, and seller can change fit, price, inventory, or safety. Do not attach a parent-level review to every variant when the reviewed property differs.
Check public and machine-readable relationships
Google's current product variant structured-data documentation describes ProductGroup, variesBy, hasVariant, and productGroupID for grouping variants in Google merchant-listing experiences. That is a Google eligibility mechanism, not proof of third-party AI recommendation.
| Identity field | Page | Feed | Schema | Marketplace | Status |
|---|---|---|---|---|---|
| Brand | Required | Required | Product brand | Seller listing | Compare |
| Product name | Stable | Title | Product name | Listing title | Compare |
| Parent group | Visible where useful | Item group ID | ProductGroup | Parent listing | Compare |
| Variant properties | Selected values | Variant attributes | variesBy/Product | Offer options | Compare |
| SKU/MPN/GTIN | Appropriate public/support source | Identifiers | Identifier properties | Platform fields | Validate |
| Seller | Offer context | Merchant account | Offer seller where used | Store | Validate |
| Condition | Visible when material | Condition | itemCondition | Listing condition | Validate |
| Canonical URL | Selected variant behavior | Link | Page identity | Product URL | Validate |
Score Public Availability and Rendering
The facts must exist in the public experience the audit intends to evaluate. A spec visible only after login, in a non-rendering widget, or inside an image is a weaker source unit.
Verify the URL and status
Record HTTP status, canonical, indexability intent, robots behavior, content rendering, mobile access, localization, and whether the selected variant survives the URL/share flow.
Inspect visible facts
Price, stock, variant, core attributes, fit, and evidence should be readable without requiring an analyst to infer them from a script object. Important gated documents may support buyers but remain unavailable to public answer retrieval.
Separate access from use
A 200 page is not proof that a source was retrieved, parsed, cited, or trusted. Award points for source-side availability and keep outcome measurement separate.
| Check | 0 points | Partial | Full |
|---|---|---|---|
| URL/status | Broken/redirect loop | Intermittent/soft issue | Stable intended response |
| Canonical | Wrong entity | Ambiguous | Correct self/declared canonical |
| Rendering | Critical facts absent | Some client-only/hidden | Critical facts public and readable |
| Variant state | Resets/wrong variant | State partly retained | Purchasable unit retained |
| Mobile | Blocked/broken | Degraded | Functional critical flow |
| Localization | Wrong market/currency | Mixed | Scope-aligned |
| Images | Image-only critical facts | Partial text alternative | Key facts in text + useful images |
| Purchase path | Dead end | Friction/ambiguity | Declared offer can be reached |
Score Category-Specific Attribute Completeness
Generic PDP templates fail when they treat every product as a name, description, price, and star rating. Buyers make category-specific decisions.
Build the attribute decision set
Start with product experts, support tickets, returns reasons, filters, fit guides, documentation, comparison criteria, and governed prompts. Record whether each attribute is required, optional, not applicable, unknown, or unsupported.
Normalize units and definitions
“Lightweight,” “compact,” and “long-lasting” are not comparable measurements. Use supported units, test conditions, tolerances, and definitions where they exist.
Preserve unknown and not applicable
Do not convert missing measurements into “no” and do not award completeness points for irrelevant fields. The denominator should include required applicable attributes only.
| Category | High-value attribute examples | Common ambiguity |
|---|---|---|
| Apparel | Size, fit, material, care, weather, model measurements | Size labels differ by market |
| Electronics | Model, storage, memory, voltage, ports, compatibility, warranty | Family feature assigned to base model |
| Beauty | Ingredients, shade, skin/hair type, allergens, use, exclusions | Formula differs by region/variant |
| Food | Ingredients, allergens, quantity, nutrition, storage, origin | Pack count/unit confused |
| Furniture | Dimensions, material, load, assembly, room fit, delivery | Packaged vs assembled dimensions |
| Auto/parts | Make, model, year, trim, position, certification | “Universal” compatibility overclaimed |
| Outdoor | Size, material, temperature/weather rating, weight, capacity | Test conditions missing |
| Subscriptions | Included products, cadence, minimum term, renewal, cancellation | Introductory price treated as ongoing |
| Attribute state | Score treatment | Buyer-facing treatment |
|---|---|---|
| Supported/current | Eligible for points | Publish value + source/condition |
| Supported but stale | Partial/zero by rule | Refresh before relying |
| Unknown | No completeness point | Label unknown; investigate |
| Not tested | No evidence point | Say not tested |
| Not applicable | Remove from denominator | Explain when confusing |
| Contradicted | Critical review | Resolve conflict |
| Unsupported claim | Zero/possible gate | Remove or narrow |
Score Fit, Constraints, and Exclusions
Product recommendations are routing decisions. A product can be strong for one use case and wrong for another.
Define best-for with conditions
Name the user, job, environment, budget, compatibility, size, material, time, and evidence. Avoid “perfect for everyone.”
Publish avoid-if statements
Surface allergies, incompatibilities, unsupported devices, size limits, climate boundaries, age restrictions, unavailable markets, installation needs, or required subscriptions where material.
Route non-fit buyers honestly
Offer a different variant, product, category, repair, rental, professional advice, or no-purchase option where appropriate. The AI-search matchmaker framework reinforces why constraints and exclusions change the recommendation.
| Constraint family | Question | Source | Failure risk |
|---|---|---|---|
| Audience | Who can use it responsibly? | Product/safety guidance | Universal recommendation |
| Use case | Which job/environment? | Specs, tests, documentation | Vague benefit claim |
| Compatibility | With which model/system? | Compatibility data | Wrong purchase/safety risk |
| Size/fit | Which dimensions/body/product? | Fit/spec tables | Returns and exclusion |
| Material/ingredients | What is present/absent? | Manufacturer/regulatory data | Allergy/preference error |
| Budget/value | Which price/unit/term? | Current offer | Misleading affordability |
| Geography | Where sold/shipped/valid? | Offer/policy | Ineligible recommendation |
| Time | Delivery, setup, use duration | Shipping/docs/tests | Impossible promise |
| Avoid-if | What condition breaks fit? | Evidence/product owner | Missing boundary |
Score Commercial Facts and Freshness
Commercial data changes faster than editorial descriptions. Price, sale periods, stock, seller, shipping, delivery, returns, warranty, subscription, and minimum quantity need independent clocks.
Match visible and submitted facts
Google Merchant Center's current product data specification requires product-data price and availability to match relevant landing-page, structured-data, and checkout facts for Google's programs. Use that consistency principle inside this audit without claiming it controls other AI systems.
Record the offer unit
State currency, tax treatment where relevant, unit, pack count, minimum order, subscription, installment, membership, sale price, and effective date. A low number without the purchasable unit is not comparable.
Use event-driven clocks
Price, availability, seller, shipping rule, returns, or warranty changes should trigger affected page/feed/schema/marketplace checks immediately, not wait for an annual content review.
Use the e-commerce AI data-freshness framework to assign canonical fact owners, define source and propagation clocks, classify conflicts, and re-test observed price, stock, variant, delivery, and policy claims without promising external refresh timing.
| Commercial field | Required context | Clock | Critical mismatch? |
|---|---|---|---|
| Price | Amount, currency, unit, eligibility | Offer update | Yes when material |
| Sale price | Original, sale, effective dates | Campaign clock | Yes when active claim wrong |
| Availability | In/out/preorder/backorder + region | Inventory clock | Yes |
| Seller | Merchant and fulfillment owner | Offer clock | Sometimes |
| Shipping | Region, cost, method, estimated window | Policy/rate clock | Material |
| Delivery | Destination-specific estimate basis | Checkout clock | Material |
| Returns | Window, condition, fees, exclusions | Policy clock | Material |
| Warranty | Provider, period, scope, exclusions | Product/policy clock | Material |
| Subscription | Term, renewal, cancellation, included unit | Billing clock | Yes |
| Quantity | Pack/minimum/unit basis | Catalog clock | Yes when price comparison changes |
Google also documents availability consistency across landing pages, structured data, checkout, and submitted product data for Merchant Center. Treat mismatches as an operational risk even when the product remains visible elsewhere.
Score Evidence, Provenance, and Reviews
Answerability requires more than a brand repeating its own claim. It also requires accurate separation of source roles.
Build a claim register
For each material claim, record the product/variant, canonical wording, source, source role, date, method, condition, boundary, and allowed short form.
Preserve review context
Record platform, reviewer type, purchase/verification status where available, date, variant, geography, sample, incentives, rating scale, and adverse themes. Do not invent reviews or present an owned testimonial as independent validation.
Keep evidence close to the claim
The claim-drift framework explains how audience, evidence, entity, time, comparison, and attribution can change across retellings. Keep conditions and boundaries in the same answer unit.
| Source role | Can support | Cannot establish alone | Confidence inputs |
|---|---|---|---|
| Manufacturer/product docs | Specification and intended use | Independent preference | Version, owner, date |
| Controlled test | Performance in stated method | Every use/environment | Method, sample, conditions |
| Certification/standard | Declared scoped compliance | Broader quality superiority | Issuer, scope, expiry |
| Verified customer review | Reported experience | Universal performance | Variant, date, incentives |
| Expert/editorial review | Tested or evaluated experience | Complete market truth | Method, independence, recency |
| Marketplace listing | Offer/review context | Manufacturer truth | Seller, listing integrity |
| Community discussion | Language and reported issues | Verified product fact | Attribution, corroboration |
| Owned testimonial | Named customer experience | Independent review | Permission, scope, context |
| Review check | Pass condition | Red flag |
|---|---|---|
| Product match | Exact product/variant clear | Family/variant conflation |
| Recency | Relevant to current version | Legacy model used as current |
| Attribution | Source/reviewer visible | Anonymous copied quote |
| Incentive | Disclosed where applicable | Hidden compensation |
| Balance | Material adverse themes retained | Cherry-picked praise only |
| Evidence | Experience separated from fact | Review claim becomes specification |
| Rating scale | Platform scale/sample visible | Ratings combined across systems |
| Permission | Quote/use permitted | Scraped/reproduced improperly |
Score Comparison and Alternative Readiness
AI shopping questions often contain several products and constraints. A product page should expose comparable facts without inventing a universal winner.
Choose decision criteria
Use category facts buyers actually compare: size, materials, performance, compatibility, price unit, availability, shipping, warranty, evidence, and exclusions.
Normalize the unit
Compare equivalent pack sizes, conditions, variants, subscription terms, currencies, taxes, and test conditions. “Cheaper” can be false when one price covers 30 units and another covers 10.
Keep alternatives valid
The alternative may be another variant, repair, rental, used item, professional solution, or no purchase. Do not disparage competitors or fill unavailable cells with assumptions.
| Comparison field | Product entry | Required qualifier |
|---|---|---|
| Entity | Exact product/variant/seller | No family conflation |
| Price | Amount/currency/unit/date | Like-for-like basis |
| Availability | Region/status/clock | Purchasable state |
| Attribute | Value/unit/method | Same definition |
| Fit | Best-for/avoid-if | Declared buyer/use case |
| Evidence | Source/date/method | Source role visible |
| Unknown | Explicit unknown | No favorable inference |
| Not comparable | Reason | Keep out of winner claim |
| Alternative | Valid next route | Honest non-fit handling |
Score Page, Feed, Schema, and Marketplace Consistency
Structured data can help a platform interpret a page, but it does not repair a false visible claim or guarantee an AI recommendation.
Treat visible content as product truth
The buyer should be able to see the important product and offer facts. Do not place a different price, rating, availability, or variant only in JSON-LD.
Use platform documentation for platform eligibility
Google's current merchant-listing structured-data guide describes how Product and Offer markup can make pages eligible for Google's merchant-listing experiences. Keep “eligible for Google presentation” separate from “recommended by AI.”
Shopify operators can apply the platform-neutral score with the Shopify GEO implementation checklist, which adds variant, metafield, collection, theme, app, feed, review, and checkout acceptance gates.
Reconcile every source at one clock
Capture PDP, JSON-LD, merchant feed, marketplace listing, checkout, PIM, and inventory source together. Record latency rather than assuming instantaneous agreement.
| Field | PDP | Schema | Feed | Checkout | Marketplace | Result |
|---|---|---|---|---|---|---|
| Product/variant | Visible selected entity | Product/ProductGroup | ID/item group | Basket line | Listing | Must align |
| Price/currency | Visible offer | Offer price | Price | Charged price | Offer | Must align by eligibility |
| Sale | Visible terms | Applicable properties | Sale + dates | Charged amount | Promotion | Align clock |
| Availability | Visible status | Offer availability | Availability | Purchasable | Listing status | Align market |
| Condition | Visible when material | itemCondition | Condition | Order line | Listing | Align |
| Seller | Offer context | Offer seller where used | Merchant | Merchant of record | Store | Align |
| Rating/review | Attributable source | Eligible aggregate/review data | Platform-dependent | N/A | Platform data | Do not combine blindly |
| Shipping/returns | Visible policy | Applicable offer details | Attributes | Final terms | Platform terms | Preserve scope |
Preserve Missing, Ambiguous, and Adverse States
A score becomes unreliable when missing facts are treated as positive, failed requests are removed, and adverse evidence is hidden.
Use a state dictionary
For every field, support present, absent, unknown, unavailable, ambiguous, stale, contradicted, not applicable, not tested, and critical-fail states.
Keep observation failure separate
A page timeout or answer-product error is not proof that the product lacks information. Retry under the collection contract and preserve both results.
Do not punish correct non-fit
A product that explicitly says “not compatible with Model X” may score better on answerability than one that omits the limitation—even if exclusion reduces recommendation coverage.
| State | Meaning | Score action | Next move |
|---|---|---|---|
| Present/supported | Current evidence resolves field | Award per rubric | Maintain clock |
| Absent | Required fact not published | Zero | Source/fill if true |
| Unknown | Owner does not know | Zero | Product/data investigation |
| Unavailable | Source could not be accessed | No silent pass | Re-observe/fix access |
| Ambiguous | Multiple plausible entities/values | Partial/zero | Normalize/clarify |
| Stale | Evidence outside accepted clock | Partial/zero | Refresh |
| Contradicted | Sources disagree | Gate review | Resolve and correct |
| Not applicable | Field does not apply | Remove denominator | Document reason |
| Not tested | Claim lacks test | No evidence point | Test or narrow |
| Adverse/non-fit | Valid limitation | Can earn boundary points | Route honestly |
Calculate Score, Completeness, and Confidence Separately
One total cannot communicate whether a product has many filled fields supported only by weak sources or a smaller set of high-confidence facts.
Calculate the raw points
For each dimension, use documented requirements. A simple model can award 0, 0.5, or 1 times the item weight for missing, partial, or complete.
Calculate evidence confidence
Assign source/entity/date/method confidence separately. Do not multiply arbitrary confidence into the public headline without showing both inputs.
Report gate status first
Use PASS, INVESTIGATE, or CRITICAL FAIL, followed by raw score, completeness, and confidence.
| Output | Synthetic formula | Interpretation |
|---|---|---|
| Raw score | Sum earned dimension points / 100 | Diagnostic coverage |
| Required-field completeness | Supported required fields / applicable required fields | Fact coverage |
| Evidence confidence | Supported weighted claims / eligible claims | Source quality/context |
| Freshness coverage | In-clock material facts / material facts | Currency of truth |
| Consistency coverage | Agreeing source-field checks / eligible checks | Cross-system agreement |
| Gate status | Worst applicable critical gate | Overrides flattering total |
Illustrative dimension scoring
The following formula is a planning device:
Dimension points = dimension weight × (complete items + 0.5 × partial items) ÷ applicable items
Do not compare scores across categories until requirements and denominators are genuinely comparable.
Work Through a Synthetic Product Example
Example TrailShell and every fact, score, source, and result below are synthetic. They demonstrate the method, not a GeoZ customer outcome.
Scope the item
The audit unit is the women's medium blue 2026 variant sold by Example Brand in the United States for a wet-weather hiking use case under $200 with delivery required within 7 days.
Keep the defects
The PDP and feed disagree on availability, the selected-variant URL resets to the parent, and a legacy review refers to the 2024 fabric. These are not edited out because the product has strong attribute content.
Apply the gate
The unresolved availability mismatch triggers INVESTIGATE, not PASS, until the offer clock and checkout state are reconciled.
| Dimension | Weight | Earned | Synthetic finding |
|---|---|---|---|
| Identity | 10 | 7 | Variant URL resets; identifiers otherwise stable |
| Public availability | 10 | 8 | Facts render; variant state weak |
| Attributes | 15 | 13 | Material, weight, size, care present; test method thin |
| Fit/constraints | 15 | 12 | Use and size clear; avoid-if incomplete |
| Commercial freshness | 15 | 8 | Feed says in stock; checkout says unavailable |
| Evidence/reviews | 15 | 9 | Evidence exists; one legacy-review mismatch |
| Comparison readiness | 10 | 7 | Units clear; alternatives incomplete |
| Data consistency | 5 | 3 | Availability and variant disagreement |
| Governance | 5 | 4 | Owners assigned; event SLA missing |
| Raw total | 100 | 71 | Partial answerability |
| Gate | Result | Decision |
|---|---|---|
| Entity | Investigate | Repair variant URL and review attribution |
| Public access | Pass | Maintain |
| Price/currency | Pass | Recheck at campaign changes |
| Availability | Investigate | Reconcile feed/PDP/checkout clock |
| Fit/safety | Pass with gap | Add avoid-if boundary |
| Evidence | Investigate | Remove or relabel legacy review |
| Final | INVESTIGATE / 71 | Not eligible for “ready” claim |
Roll Up the Catalog Without Hiding Risk
The arithmetic mean is a poor portfolio summary when critical errors cluster in high-revenue, regulated, or frequently recommended products.
Show the distribution
Report count and share in score bands, critical fails, unresolved variants, and missing confidence states.
Add business exposure separately
Revenue, traffic, inventory value, margin, return rate, or strategic priority can help order repairs. They do not change whether a fact is true.
Keep categories separate
An attribute rubric for cosmetics cannot be compared directly with one for electronics until the category requirements and severity model are normalized.
| Portfolio metric | Synthetic result | Why the mean is insufficient |
|---|---|---|
| Products audited | 120 | Defines population |
| Variants audited | 480 | Shows purchasable-unit scale |
| Mean raw score | 78 | Can hide severe tails |
| Median raw score | 82 | Distribution still needed |
| 90–100 | 22 products | Strong group |
| 75–89 | 51 products | Material gaps remain |
| 60–74 | 29 products | Repair priority |
| Below 60 | 18 products | Weak foundation |
| Critical fails | 14 products | Must remain visible |
| High-revenue critical fails | 6 products | Priority input, not truth modifier |
| Unresolved variant conflicts | 37 variants | Entity risk |
| Stale commercial facts | 64 variants | Freshness risk |
The following product-level roll-up is synthetic and exists only to demonstrate how raw scores, confidence, variants, and gate failures remain separate.
| Product ID | Variants audited | Raw score / 100 | Evidence confidence % | Critical fails | Priority / 5 |
|---|---|---|---|---|---|
| EX-P01 | 12 | 94 | 92 | 0 | 2.1 |
| EX-P02 | 8 | 88 | 81 | 0 | 2.8 |
| EX-P03 | 16 | 84 | 76 | 1 | 4.7 |
| EX-P04 | 4 | 79 | 90 | 0 | 3.0 |
| EX-P05 | 24 | 76 | 68 | 2 | 4.8 |
| EX-P06 | 6 | 72 | 73 | 0 | 3.6 |
| EX-P07 | 18 | 69 | 61 | 1 | 4.4 |
| EX-P08 | 3 | 63 | 85 | 0 | 3.2 |
| EX-P09 | 20 | 57 | 54 | 3 | 5.0 |
| EX-P10 | 9 | 41 | 48 | 2 | 4.9 |
Connect the Score to a Prompt Panel
The score audits source readiness. A prompt panel observes answer behavior. Use both without collapsing them.
Select eligible products and routes
Include a representative set of categories, constraints, markets, variants, and business priorities. Do not choose only high-scoring products or prompts that name the brand.
Keep the observation contract
Record prompt, route, product eligibility, answer product/mode, market, language, time, repeats, role state, citations, accuracy, fit, and missing output. The 50-query evaluation-panel guide provides the broader governance model.
Compare diagnosis, not just totals
A high-answerability product that is absent suggests a different investigation from a low-answerability product that appears with an inaccurate price.
| Answerability | Observed answer | First investigation |
|---|---|---|
| High | Absent | Eligibility/retrieval/source competition |
| High | Mentioned, not compared | Decision-route/competitive evidence |
| High | Recommended accurately | Maintain; re-observe variance |
| High | Recommended inaccurately | Third-party drift/entity/source review |
| Low | Absent | Repair source truth before conclusions |
| Low | Mentioned inaccurately | Identity/freshness/claim correction |
| Low | Recommended | Risk review; do not celebrate blindly |
| Critical fail | Any | Correct material issue before optimization |
The E-GEO paper explainer is useful for understanding intent-rich queries, candidate retrieval, LLM re-ranking, factuality, and optimization loops. Preserve its limitation: the described benchmark used a simulated generative-shopping setup and rank-improvement objective, not a universal production recommendation system.
Turn Findings Into an Action Queue
Every gap needs a failure-layer code and owner. “Write more content” is not an acceptable default.
Separate repair types
Use fix entity, fix page/rendering, fix PIM/feed, fix commercial operations, add true attribute, add evidence, correct third party, change product/policy, create page, observe, or no action.
Score urgency independently
Combine materiality, business exposure, buyer-route importance, evidence readiness, severity, and effort transparently. A critical safety mismatch can outrank a high-traffic content opportunity.
Keep the product decision open
If the item genuinely lacks a required capability, the right outcome may be to exclude it from that route or improve the product—not produce persuasive copy.
| Failure layer | Example | Owner | Action |
|---|---|---|---|
| Entity | Variant names conflict | PIM/merchandising | Normalize IDs/names/URLs |
| Technical | Core facts not rendered | Engineering/SEO | Fix public delivery |
| Commercial | Price/stock mismatch | Ecommerce ops | Reconcile sources/clocks |
| Attribute | Required dimension absent | Product/data | Source and publish true value |
| Fit | Avoid-if missing | Product/content | Add bounded routing |
| Evidence | Claim has no method | Product/research | Test, source, or narrow |
| Review | Legacy variant credited | CX/content/legal | Correct attribution |
| Comparison | Units incompatible | Merchandising/content | Normalize criteria |
| Product | Required compatibility absent | Product | Exclude or change product |
| Observation | Answer unavailable | Measurement | Re-run; preserve missingness |
Illustrative priority queue
| Item | Severity 1–5 | Exposure 1–5 | Readiness 1–5 | Effort 1–5 | Synthetic priority |
|---|---|---|---|---|---|
| Wrong allergen statement | 5 | 5 | 5 | 2 | 4.75 |
| Variant price mismatch | 5 | 4 | 5 | 2 | 4.45 |
| In-stock/checkout conflict | 5 | 4 | 4 | 2 | 4.25 |
| Compatibility ambiguity | 5 | 3 | 3 | 3 | 3.55 |
| Missing comparison unit | 3 | 4 | 4 | 2 | 3.55 |
| Weak review provenance | 3 | 3 | 3 | 3 | 3.00 |
| Generic intro copy | 1 | 2 | 5 | 2 | 2.05 |
Every score and weight in this table is illustrative.
Run a 30/60/90-Day Program
A 90-day program can establish scope, audit a prioritized catalog, correct critical defects, and create a repeatable observation loop. It cannot guarantee AI visibility or sales inside 90 days.
Days 1–30: define and gate
Choose categories, products, variants, markets, prompts, material attributes, source roles, clocks, and gate rules. Audit the highest-exposure products and stop unsafe or materially wrong claims.
Days 31–60: repair systems and evidence
Fix entity relationships, rendering, PIM/feed/schema/page mismatches, commercial clocks, attribute gaps, evidence, and review attribution. Create new content only for a distinct buyer decision.
Days 61–90: observe and institutionalize
Run the fixed prompt panel, code answer states, compare diagnosis to the audit, build the change register, and assign event-driven SLAs.
| Phase | Days | Synthetic output | Gate |
|---|---|---|---|
| Scope | 1–5 | 2 categories, 20 products, 80 variants | Owners approve unit |
| Requirements | 6–10 | 60 applicable attribute rules | Product/legal review |
| Critical audit | 11–18 | 20 products × 8 gates | Material errors contained |
| Scoring | 19–25 | 20 scorecards + confidence | Evidence sampled |
| Queue | 26–30 | 35 actions, 12 owners | Capacity approved |
| Entity/data repair | 31–42 | 18 conflicts resolved | Source systems agree |
| Commercial repair | 43–50 | 24 offers reconciled | Page/feed/checkout align |
| Evidence/content | 51–60 | 10 answer units, 4 evidence updates | Claims pass review |
| Panel | 61–72 | 50 prompts × eligible products | Collection contract passes |
| Diagnosis | 73–82 | Role/accuracy/fit matrix | Missingness retained |
| Governance | 83–90 | 1 dashboard, clocks, next queue | Continue/investigate/act/defer |
Capacity checklist
- 1 product taxonomy owner approves parent, product, and variant relationships.
- 1 ecommerce-operations owner governs price, availability, shipping, and returns clocks.
- 1 SEO/GEO owner governs public rendering, canonicals, prompts, and observations.
- 1 product/content owner governs attributes, fit, exclusions, and comparisons.
- 1 evidence owner governs test methods, certifications, reviews, and source roles.
- 2 category rubrics remain separate until their requirements are normalized.
- 20 products and 80 variants form the synthetic pilot—not a universal sample size.
- 8 critical gates are reviewed before any raw-score grade is used.
- 9 dimensions add to 100 points under rubric version PAS-1.0.
- 10 missing/adverse states remain available to coders.
- 50 prompts cover declared routes rather than all private demand.
- 3 answer products with 1 eligible mode each form the illustrative panel.
- 2 runs per prompt provide observations, not statistical certainty.
- 5 source roles are sampled for high-risk claims.
- 0 fabricated reviews, attributes, specifications, certifications, or offers are allowed.
- 0 Direct sales or revenue are attributed from the score alone.
- 4 executive dispositions remain: continue, investigate, act, or defer.
- 1 change register preserves every corrected material fact and affected URL.
What GeoZ Delivers in a Product Answerability Audit
GeoZ is a Value as a Service company for SEO and GEO. A product-answerability audit turns catalog truth, public evidence, and answer observations into a prioritized action queue.
Product and variant map
GeoZ can map declared products, variants, identifiers, sellers, markets, pages, feeds, schema, marketplaces, and material decision attributes inside the agreed scope.
Evidence and scorecard
The work can apply critical gates, the transparent rubric, evidence confidence, freshness, consistency, and missingness without presenting the result as an AI-ranking predictor.
Measurement-to-execution loop
The audit can connect source readiness to a governed prompt panel, observed answer roles, accuracy, fit, corrective actions, and re-observation. The How GeoZ Works guide explains the broader operating model.
| Work package | Inputs | Outputs | Explicit boundary |
|---|---|---|---|
| Scope/entity map | Catalog, variants, sellers, markets | Governed audit units | No entity inference without review |
| Requirements | Category experts, support, prompts | Applicable attribute/constraint rubric | Not all demand |
| Critical gates | Page/feed/checkout/evidence | Pass/investigate/fail register | No average override |
| Scorecard | Nine dimensions | Raw score + confidence + freshness | Not ranking predictor |
| Consistency audit | PDP, schema, feed, marketplace | Field-level conflicts | Platform eligibility kept separate |
| Evidence audit | Claims, tests, reviews, certificates | Source roles and gaps | No fabricated corroboration |
| Prompt observation | Eligible products/routes/modes | Role, source, accuracy, fit states | No private buyer count |
| Action queue | Severity, exposure, readiness, effort | Owners and next actions | No guaranteed visibility/sales |
If your catalog has rich PDPs but AI answers still conflate variants, repeat stale prices, or omit critical fit information, request a product answerability audit. Bring the category scope, top products, feed/PIM owner, material constraints, current schema, and known commercial-data defects.
Use the Final Review Gate
The scorecard should not ship until product, ecommerce operations, evidence, technical delivery, and measurement owners agree on the scope and unresolved defects.
Review product truth
Confirm identity, attributes, compatibility, safety, fit, commercial terms, evidence, reviews, alternatives, and limitations.
Review platform-specific statements
State that Google documentation governs Google eligibility and merchant data. Do not imply that Product schema or Merchant Center approval guarantees inclusion in ChatGPT, Gemini, Perplexity, Claude, Meta AI, or any other answer surface.
Review measurement language
Use the GeoZ Metrics Dictionary to keep answerability, eligibility, presence, citation, comparison, recommendation, referral, sale, revenue, and causality separate.
| Final gate | Pass condition | Failure response |
|---|---|---|
| Scope | Purchasable unit and buyer decision fixed | Rescope |
| Entity | Product/variant/seller consistent | Fix taxonomy/IDs |
| Access | Important facts publicly render | Fix delivery |
| Attributes | Required applicable facts supported | Source or mark unknown |
| Fit | Best-for/avoid-if/alternative visible | Add boundaries |
| Commercial | Price/stock/shipping/returns current | Reconcile systems |
| Evidence | Claims attributable and in scope | Narrow/remove/source |
| Consistency | Page/schema/feed/checkout aligned | Correct source of truth |
| Gates | No unresolved critical fail hidden | Investigate before pass |
| Measurement | Score not called visibility predictor | Rewrite claim |
The GEO Community's Meta AI shopping research explainer is a useful source for thinking about product facts, offers, reviews, comparisons, structured data, and interface testing. This article does not reuse its platform-specific adoption, recommendation-count, checkout, or earned-media figures as universal facts.
Key Takeaways
Score answerability, not popularity
The rubric measures whether a defined product question can be answered from public, current, attributable sources. It does not predict an answer system's ranking or recommendation.
Let critical facts override averages
Variant identity, commercial truth, compatibility, safety, and evidence can invalidate a flattering raw score. Keep gate status first.
Preserve source roles and clocks
PDP, schema, feed, checkout, marketplace, manufacturer, review, test, and observed-answer data have different jobs. Reconcile them without pretending they are independent or simultaneous.
Route each gap to the correct owner
Fix data when data is wrong, evidence when evidence is weak, rendering when facts are hidden, product when the capability is absent, and content only when a distinct answer unit is truly missing.
FAQs
What is an e-commerce Product Answerability Score?
It is a transparent self-assessment of whether a defined product or variant can be identified, accessed, compared, qualified against constraints, represented with current commercial facts, and supported by evidence. The rubric in this article is illustrative and is not a validated AI-ranking or recommendation predictor.
Does Product schema improve the answerability score?
Consistent Product and Offer markup can contribute to the page/feed/schema consistency dimension and can support eligibility for documented Google merchant experiences. It does not repair wrong visible facts, replace category-specific content, or guarantee retrieval or recommendation in any AI product.
Should we score every SKU or only product pages?
Score the purchasable unit whenever variants materially change fit, price, inventory, safety, compatibility, seller, condition, or evidence. Start with a risk- and business-prioritized sample, then expand once the rubric, source clocks, and owners are stable.
What happens when a product has a critical fail but a high raw score?
Report the raw score for diagnosis, but set the final gate status to CRITICAL FAIL or INVESTIGATE. Do not call the product ready until the material identity, access, commercial, compatibility, safety, or evidence problem is resolved and rechecked.
Can reviews make a product more answerable?
Relevant, current, attributable reviews can support experience, fit, durability, usability, and adverse-theme evidence. Keep product/variant match, incentives, method, date, platform, rating scale, and source independence visible. Never fabricate or decontextualize reviews.
How do we know whether improving answerability changed AI recommendations?
Version the product sources and run a governed prompt panel before and after the accepted change while preserving answer product, mode, market, clocks, eligibility, repeats, missingness, role states, accuracy, and competing events. Movement is an observation unless the design supports a stronger causal claim.