Marketplace systems / Game discovery

Game Discovery Is an Evidence Problem

A game-marketplace application profile for source-qualified discovery claims.

Jason Doyle 28 September 2026 42 minute read

Disclosure: These views are my own and do not represent any current or former employer. This paper uses public platform documentation, technical standards and published research. It does not describe non-public Steam architecture, ranking models or commercial plans.

Artifact status: This draft includes a machine-validated JSON Schema, a synthetic conformance example and a reproducible five-game Steam case study. The case study contains public source snapshots, 15 human-checked review excerpts, ten discovery queries and four adversarial fixtures. It is not a production marketplace integration, player study or representative benchmark.

Executive summary

A player can already ask an agent to find a product, compare offers and complete a purchase. Meta launched Muse in September 2026 with browser-based shopping and payment support. Amazon then blocked Muse from shopping on Amazon.com, while Shopify enabled Meta as an agentic sales channel and documented Shop Pay as a Muse wallet or checkout option when available.[1][2][3]

The disagreement is partly about control over the customer interface. A store usually decides how products are searched, ranked and presented. An external agent can interpret the user's request before the store sees it. It may query several sellers, remove candidates that fail the user's constraints and apply its own final ordering.

The underlying commerce stack is already substantial. The Universal Commerce Protocol supports catalogue search and product lookup against a known business, along with checkout and order capabilities. The Agentic Commerce Protocol supports product feeds, seller capability discovery and delegated checkout. AP2 provides signed evidence that a user authorised a checkout and payment, including human-not-present flows. Visa and Mastercard provide additional agent identity, catalogue and payment services.[6][7][8][9][10][11]

This means the open problem is narrower than a new agent-commerce protocol.

Games create a difficult discovery problem because the information needed for a useful choice comes from sources with different authority. A marketplace can publish a developer's controller declaration or a compatibility result from its own test. A developer can state that a campaign supports four players. Reviews can describe coordination burden or frustration. Marketplace telemetry can measure session distributions for a defined population. A model can infer a claim from text. These records should not collapse into one unqualified field.

Consider a request such as:

Find a cooperative game for three people this evening. It must support our target platforms, avoid a second launcher and work in sessions under an hour. At least one of us should not already own it.

The request contains several kinds of constraint:

  • target-platform and multiplayer support are product facts;
  • launcher requirements belong to a specific offer or edition;
  • ownership is private account state;
  • session fit is a population claim that depends on mode and cohort;
  • the final ordering is a recommendation decision.

Current public game metadata can represent some of this. Schema.org includes VideoGame, SoftwareApplication, Product, Offer and Review. The Video Game Metadata Schema covers target-platform editions, gameplay features and version information. Steam exposes store metadata, tags and reviews, while its own discovery surfaces use several distinct systems.[12][14][15][25][26][46]

The proposed integration layer is a game-domain application profile that preserves how each claim was produced and where it applies.

This paper's claim is:

Agent-mediated game discovery needs a game-domain application profile that binds each discovery claim to its identity, source, method and applicability. The profile should preserve unknowns while exposing only scoped eligibility and placement context.

This paper proposes a Game Discovery Profile with five boundaries:

  1. It distinguishes the game work, target-platform edition, release and commercial offer.
  2. It separates functional facts, marketplace verification, aggregate observations, population experiences, model inferences and personalised predictions.
  3. It records source, method and applicability for each claim, and an observation window for aggregate evidence.
  4. It lets a marketplace answer private eligibility questions without returning a user's library or social graph.
  5. It reports ranking context and commercial influence without requiring the marketplace to publish its ranking weights.

The profile is an application design built on existing provenance standards. W3C PROV already models entities, activities, agents and derivation. Web Annotation can identify the exact source fragment supporting a claim. SHACL can validate graph constraints. The proposed JSON Schema uses a smaller serialisation suitable for a prototype and leaves a standards-aligned RDF mapping as later work.[35][36][37]

The paper also avoids claiming a new high-level recommendation architecture. Personal Agent-Mediated Recommendation already argues for agents that gather distributed evidence while preserving source identity, disagreement, freshness and privacy boundaries. GAVEL already uses the term Evidence Contract for claim-level fact-checking. The contribution here is the game-marketplace profile, its public schema and a reproducible evaluation method.[33][34]

The public case study generated validated profiles for Portal 2, Deep Rock Galactic, Hades, Balatro and Left 4 Dead 2 from public Steam data. It retained the exact review query parameters, source-response hashes and short evidence excerpts without Steam account identifiers or recommendation IDs. The excerpts and fingerprints remain linkable to the underlying public reviews and are not anonymised.

A flattened baseline and a profile-aware engine agreed on all ten clean queries, returning the same top result on the nine that produced a candidate. One query returned no candidate from either engine because the selected public data did not establish whether any offer avoided an external launcher. Both engines preserved that unknown instead of inferring a value.

Four adversarial fixtures then added a model-inferred co-op claim, a release-mismatched claim, duplicated review evidence and a prompt-injection string. The flattened baseline promoted the targeted game in three fixtures and returned one top result that violated a hard constraint. The profile-aware engine recorded no target promotion and no hard-constraint violation.

These results demonstrate the behaviour of the contract checks in a small, deterministic case study. They do not establish recommendation quality for players or resistance to adaptive attacks.

The marketplace remains important in this model. It owns release identity, private eligibility checks and abuse controls. It may generate the candidate set and retain private behavioural signals. The agent interprets the user's request and may rerank candidates under the authority the user granted.

Agent-mediated discovery changes the interface between user intent, marketplace evidence and ranking while leaving the store in place.

1. Agentic commerce has reached the catalogue

1.1 The interface is now contested

Muse was announced as a personal agent able to navigate websites and complete multi-step tasks. Meta said credentials were held in secure storage outside the model and that Shop Pay support would follow its initial payment integration.[1]

Amazon blocked Muse shopping later that month. Its public position was that third-party purchasing applications should identify themselves and respect a service provider's decision about participation. Amazon's Conditions of Use also require agents to identify each request and prohibit attempts to evade access controls.[2][45]

Shopify took a different route. It added Meta as an AI channel, made product sharing available through Shopify Catalog and gave merchants controls over catalogue access and direct checkout. Its published help documentation limits availability by store eligibility, geography and product type. Merchants can disable the channel or direct checkout.[3]

These choices show that agent access is a platform policy question. The store may permit an agent, provide a managed interface or refuse access. Participation does not follow automatically from a user asking an agent to shop.

The same month, six banks published voluntary principles for trusted agentic commerce. They called for clear agent identification, transparency around prioritisation and disclosure of sponsored options. Reuters reported those principles alongside John Lewis data showing that searches attributed to agents had risen from 0.3 per cent to 2.5 per cent over one year. The figure describes searches at one retailer. It is not a marketwide purchase share.[4][5]

1.2 Existing protocols cover different layers

Agentic commerce is a stack with several interfaces.

Layer Public mechanism Main responsibility
Catalogue and capability discovery UCP, ACP, managed commerce networks Find products and learn what a known seller supports
Checkout orchestration UCP and ACP Create carts, calculate terms and place orders
User authorisation AP2 and related mandate systems Prove what the user allowed the agent to buy
Agent recognition HTTP signatures, Visa TAP, network registration Distinguish an approved agent from unidentified automation
Payment-credential authorisation and scoping AP2 and card-network agent payment programmes Bind payment authority to an approved transaction

UCP's August 2026 release defines catalogue search and lookup against a business's catalogue. Its .well-known profile document is fetched from a business domain that is already known. It also separates namespace provenance from trust: controlling a schema namespace proves who published it, not whether its claims are correct.[6]

ACP supports product feeds and a seller discovery document. An agent platform can ingest merchant products into its own catalogue service. ACP also carries affiliate attribution, while leaving weighting and settlement to the participants.[7]

AP2 works at a different boundary. It binds checkout and payment mandates to user authority. Catalogue APIs and the method used to infer the user's task are outside its scope.[8]

Visa's Trusted Agent Protocol lets a merchant recognise and verify an approved agent, including one initially unknown to that merchant. Visa Intelligent Commerce Connect is a protocol- and token-vault-agnostic on-ramp that also makes merchant catalogues discoverable. Mastercard Agent Connect offers merchant-approved catalogue access and cart orchestration through a managed network.[9][10][11]

The whitepaper opportunity therefore starts after generic product discovery exists. A game marketplace needs to decide which game-specific records an agent may query and how those records preserve evidence.

1.3 Discovery and ranking are separate operations

Finding a candidate does not determine where it appears.

A marketplace can:

  • reject an ineligible game before ranking;
  • generate a candidate set using private behavioural signals;
  • attach public evidence records;
  • mark a commercial placement;
  • allow the agent to rerank the result.

An agent can:

  • translate natural language into constraints;
  • decide which marketplaces receive which parts of the request;
  • combine compatible candidate sets;
  • apply user-specific preferences;
  • explain its final choice.

The marketplace may still own most of the discovery machinery. The external agent changes the final interface and can add another ranking layer.

This distinction matters because catalogue interoperability does not require publishing a proprietary recommender. A platform can return eligible candidates and source-qualified claims while retaining model weights, embeddings and private experiment data.

2. A game offer is more than a product record

2.1 Work, edition, release and offer are different identities

One game can exist as several target-platform editions. An edition can have multiple packages, regional offers and bundles. A live game can change materially without changing its store identity.

An agent answering "do we already own this?" needs more than a title string. It must distinguish:

creative work
  -> target-platform edition
    -> release or build
      -> package or bundle
        -> regional offer
          -> account entitlement

The same title may be:

  • included in a subscription on one storefront;
  • owned as a base game without required DLC;
  • available through a key that activates elsewhere;
  • blocked in a region;
  • incompatible with the player's hardware;
  • duplicated inside a bundle.

Schema.org can represent games, software, products and offers. VGMS provides more detailed game metadata, including target-platform editions and version information. The Game Ontology Project provides an analytical vocabulary for rules, goals and game entities.[25][26][27][46]

These systems establish substantial prior art. The profile proposed here does not replace them. It connects identity and offer records to evidence used for agent-mediated discovery.

2.2 Functional facts and experiential claims need different treatment

Some questions have a release-specific answer:

Does this edition support Windows?
Does the current build support online co-op?
Is kernel-level anti-cheat present?
Does the offer require another launcher?

Other questions ask about an experience:

Can a group make progress in forty-minute sessions?
How much coordination does co-op require?
Is returning after several months difficult?
Does progression create recurring pressure to play?

Experiential properties depend on the player and context. Challenge can be performative, emotional, cognitive or decision-making. PXI separates functional qualities from psychosocial outcomes. Existing work on game challenge also shows why one difficulty value is inadequate.[28][29]

The profile should therefore distinguish:

Claim class Example Suitable evidence
Functional fact Supports online co-op for four players Release metadata and marketplace verification
Marketplace verification Passes a defined compatibility test Test profile, release identity and result
Aggregate observation Median session duration for a cohort Thresholded marketplace telemetry with a time window
Population experience Players report high coordination burden Defined study or annotated reports with uncertainty
Model inference Reviews suggest frequent interruption points Source-linked extraction with a method version
Personalised prediction Likely to fit this user's evening sessions User constraints, calibrated model and supporting claims

The classes can disagree without one record being malformed. A developer may describe drop-in play as low-friction while player reports identify difficult role handovers. The response should preserve both sources and mark the claim as contested.

2.3 Time changes the answer

A game can change after:

  • a balance patch;
  • a new matchmaking mode;
  • a server-region closure;
  • an accessibility update;
  • an anti-cheat change;
  • a progression redesign.

Reviews written before that change remain part of the public record. Their applicability may be lower for the current release.

Steam's review API exposes creation and update timestamps, reviewer playtime, whether the reviewer purchased the game on Steam and Early Access status. It also lets callers include or exclude periods marked as off-topic activity. A reproducible aggregate must record those query choices.[18][19]

A label such as 90% positive omits:

  • the observation window;
  • the language filter;
  • the purchase filter;
  • the off-topic setting;
  • the game version;
  • the numerator and denominator.

The profile treats retrieval time, observation time and validity as separate fields. A new release can invalidate an old claim without deleting its history.

2.4 Private state can be evaluated without being exported

Ownership, wishlists, playtime and friend relationships are account data. Steam groups purchased and wishlisted games, achievements and playtime under its game-details privacy setting.[22]

An agent may need an answer such as:

At least one member of this authorised group does not own the required edition.

It does not need every library entry.

A marketplace can evaluate the constraint and return:

{
  "eligibility_status": "all_eligible",
  "ownership_status": "partially_owned",
  "disclosure": "aggregate_only",
  "authorisation_ref": "synthetic-authorisation-2026-09-28-001",
  "group_size": 3,
  "checked_at": "2026-09-28T15:59:00Z",
  "expires_at": "2026-09-28T16:09:00Z"
}

eligibility_status reports whether the authorised participants can use the offer. ownership_status reports whether none, some or all of them already own the required edition. Both fields also accept unknown, meaning the marketplace could not complete the check, and not_evaluated, meaning the check was outside the request.

The result reveals enough for the current decision. It does not publish which person owns which title.

3. Steam's public evidence already has different owners

3.1 Tags mix developer, player and moderator input

Steam tags are applied by developers, non-limited player accounts and Steam moderators. Developers can order tags in the Tag Wizard. As players apply tags, the relative weights of those tags can change over time. The top 20 tags influence visibility and several tag-driven recommendation surfaces.[12]

This makes tags useful collective metadata. It also removes source separation from the public aggregate. A consumer sees the weighted result without the individual contribution from each source class.

Valve's May 2026 tag revision removed Well-Written and Masterpiece because they were subjective and inconsistently applied. It also removed franchise-specific tags because community tagging was a poor source of official identity.[13]

Those removals illustrate two boundaries:

  • some experiential judgements are too unstable for one catalogue label;
  • some facts need an authoritative source.

An agent-facing profile should not recreate either problem under new field names.

3.2 Steam uses several discovery systems

Steam's public documentation does not describe one platform-wide ranking score.

Its visibility documentation names purchases and play as strong signals. It also says language support and accurate tags matter. Store-page traffic and conversion rate do not directly determine visibility. Review score does not affect visibility while it remains at Mixed or above. A score below 40 per cent reduces the likelihood of featuring.[14]

Valve's 2019 and 2020 descriptions of the Interactive Recommender said it used a neural model trained on playtime patterns from millions of players and billions of sessions. Tags and review scores did not feed that described core model. Release date was the game-side input Valve identified, with tags available as a later filter.[15][16]

Steam's personalised tag and category hubs use signals including play history, network of friends, followed developers and wishlist state.[17]

Claims about one surface should not be projected onto another. A public tag relationship does not reveal the Interactive Recommender. A playtime model does not describe the tag and category hubs.[17]

The proposed profile sits below that private ranking layer. It describes candidate facts and evidence that a marketplace may choose to expose.

3.3 Reviews carry useful attribution and known weaknesses

Steam reviews expose more context than a binary score. Public records can include:

  • reviewer playtime;
  • whether the copy was purchased on Steam;
  • free-copy and Early Access status;
  • update timestamps;
  • helpfulness votes;
  • an attributed developer response.

Valve separates review visibility from score calculation. Key-activation reviews remain readable but do not count towards the score. Periods classified as off-topic review bombing remain visible while being excluded from the default aggregate. A newer helpfulness system changes ordering without changing the score.[18][20][21]

These controls show why the original source and transformation method matter. A model that scrapes the first visible reviews receives a result shaped by ordering and moderation policy. A model that calls the review API receives a result shaped by its filters.

Research on Steam reviews found that positive and negative reviews differed in playtime patterns. An attempted LDA analysis did not produce meaningful topics, possibly because reviews used extensive game-specific terminology. Reviews were useful, but they were not an objective record of player experience.[30]

3.4 Verification should remain visible

Steam Deck compatibility is based on a Valve review against published criteria. Controller and accessibility features are primarily declared by developers through structured interfaces.[23][24][47]

Both can appear as catalogue fields. Their evidence differs.

Public field Main source What the field establishes
Steam Deck compatibility Platform test Result against a defined hardware and review profile
Controller support Developer declaration Claimed support for named controller categories
Accessibility feature Developer declaration Claimed presence of a structured feature
User tag Mixed community aggregate Weighted description or judgement
Review score Filtered player aggregate Sentiment under the marketplace's score policy
Current player count Marketplace measurement Connected players at one point in time, excluding disconnected play [48]

An agent can make a better explanation when these distinctions survive the API boundary.

4. Agent mediation changes the trust boundary

4.1 One request can be split across several systems

A discovery flow may look like this:

player
  -> personal agent
    -> marketplace A candidate service
    -> marketplace B candidate service
    -> entitlement checks
    -> evidence records
  -> agent-side constraint evaluation
  -> final shortlist
  -> marketplace checkout

The agent does not need the same data at every step.

Candidate generation may require only broad constraints. Eligibility checks may need account authority. Final explanation may need source-qualified evidence. Payment requires a separate mandate.

The user should be able to grant these operations independently.

4.2 The marketplace retains important controls

The marketplace should control:

  • offer and release identity;
  • private account evaluation;
  • candidate eligibility;
  • access to aggregate telemetry;
  • abuse detection;
  • rate limits and commercial terms.

The agent should control:

  • interpretation of the user's request;
  • disclosure of user preferences to each marketplace;
  • cross-marketplace comparison;
  • final reranking under user instructions;
  • the user-facing explanation.

Some deployments will combine these roles. A marketplace may provide its own agent. The boundary still matters because it identifies which evidence and authority each function uses.

4.3 A recommendation explanation has two parts

An evidential explanation answers:

What supports the claim that this game fits short cooperative sessions?

A decision explanation answers:

Did that claim materially affect the ranking?

The first can be correct while the second is misleading. A generated explanation may cite session evidence even if popularity dominated the score. Research on recommender explanations distinguishes several user-facing purposes, including transparency and effectiveness. Separate work asks whether an explanation reflects the decision process that produced the recommendation.[42][49]

The profile can supply evidence references and reason codes. A faithful decision explanation also needs a logged scoring trace or a counterfactual test.

4.4 Commercial influence should be machine-readable

The six-bank principles call for transparency when prioritisation includes sponsored options. UCP transports attribution parameters without defining an attribution model. ACP carries affiliate fields while leaving weighting and settlement outside the protocol.[5][6][7]

The profile therefore includes placement values:

organic
sponsored
affiliate
unknown

This does not reveal the ranking algorithm. It states a commercial relationship relevant to the recommendation.

5. The Game Discovery Profile binds claims to identity and source

5.1 The profile reuses existing standards

The proposal reuses existing layers:

  • UCP or ACP for catalogue transport and commerce;
  • schema.org or VGMS for descriptive game metadata;
  • W3C PROV concepts for source and derivation;
  • platform APIs for verified and private state;
  • JSON Schema for the current prototype validation.

The profile defines which fields are required for this use case and how their meanings relate. It applies universal standards to one domain.

Personal Agent-Mediated Recommendation is the closest high-level prior art. It argues for source-aware evidence collection, privacy-budgeted disclosure and preservation of uncertainty. Its partial proof of concept uses 200 constructed hard Yelp restaurant-recommendation tasks. It does not define a game-domain interchange profile.[33]

GAVEL binds atomic claims to evidence for fact-checking and uses deterministic validation. Its Evidence Contract terminology predates this paper. The Game Discovery Profile uses a different name and applies claim binding to marketplace discovery.[34]

5.2 The profile binds four identities

The record begins with:

{
  "$schema": "https://jasondoyle.ie/whitepapers/game-discovery-is-an-evidence-problem/game-discovery-profile.schema.json",
  "issuer": {
    "id": "example-store",
    "role": "marketplace"
  },
  "game_identity": {
    "canonical_work_id": "urn:example:game:northstar-relay",
    "edition_id": "urn:example:game:northstar-relay:standard-pc",
    "storefront_id": "example-store",
    "release_id": "northstar-relay-1.4.2"
  },
  "offer": {
    "seller_id": "example-store",
    "offer_id": "example-1001-standard-us",
    "region": "US",
    "availability": "available",
    "requires_external_launcher_status": "known",
    "requires_external_launcher": false,
    "grants_edition_ids": [
      "urn:example:game:northstar-relay:standard-pc"
    ]
  }
}

The canonical work supports cross-store comparison. The edition distinguishes target-platform packaging and features. The release bounds technical and experiential claims. The offer carries price, region, granted editions and seller-specific requirements.

Offer-level booleans use an accompanying status field. An absent launcher fact is represented as unknown, so it is never interpreted as false.

A production system needs governance for canonical identifiers. The draft schema does not solve that registry problem.

5.3 Claims remain source-qualified

Each claim contains:

claim identifier
asserting party
property
claim class
value status
value
applicability
evidence
uncertainty
privacy mode
status

Each record names its issuer. A claim added by an agent carries its own asserting party, so a marketplace is not credited with an inference it did not supply.

The evidence object records source type, retrieval time and method. Aggregate records require an observation window and sample size, and can add a filter set.

W3C PROV supplies a richer model for entities, activities, agents and derivation. Web Annotation can point to a quoted fragment or table cell. The draft JSON shape keeps those ideas small enough for a working prototype and can carry source URIs and snapshot digests.[35][36]

5.4 Applicability is part of the claim

The sentence "median session length is 46 minutes" is incomplete.

The example record qualifies it with:

release: 1.4.2
mode: cooperative campaign
regions: US and CA
cohort: players with at least three completed sessions
window: 20 to 27 September 2026
idle timeout: 15 minutes
sample size: 12,480 users

Another region or mode may produce a different distribution. New players may have longer sessions than experienced groups.

The profile stores distributions where possible. A median and quartiles are more useful than a label such as short.

5.5 Uncertainty is decomposed

One confidence number can hide several problems.

The schema separates:

  • sampling uncertainty;
  • model uncertainty;
  • extraction uncertainty;
  • source disagreement;
  • temporal uncertainty;
  • applicability uncertainty.

Each uncertainty value needs a stated interpretation. A model probability should name its calibration method. A sample estimate should carry its interval and population. A disagreement record should preserve the conflicting sources.

Fresh provenance does not prove a claim is correct. It only makes the source and process inspectable.

5.6 Status preserves correction history

A claim can be:

active
contested
stale
withdrawn

New evidence can supersede an earlier claim without erasing it. A major release can mark observations stale. A marketplace can withdraw a verification result after re-testing.

This supports audit and avoids treating one mutable catalogue value as the complete historical record.

The draft schema accepts versioned 0.x profiles and provides x_extensions objects at record, claim and evidence level. An extension still needs a governing namespace and conformance rules before independent implementations can rely on it.

5.7 The local artifacts separate conformance and evidence

The easiest way to inspect the complete artifact set is to download and extract game-discovery-evidence-bundle.zip. The bundle SHA-256 digest is 64e7a2d52d76092c463da911a5d8d7c5fdf9e80718411560aeea20ee757bd431. It contains a README, file-level licensing, a payload manifest and the artifacts listed below.

The Game Discovery Profile JSON Schema defines the record shape. A synthetic conformance example demonstrates identity, a private aggregate entitlement result and three claim classes. It contains no real game, user or marketplace data.

The public case study is separate:

Artifact Purpose
Artifact bundle Deterministic ZIP containing the complete reproducibility package
Human annotations Selected review fingerprints, excerpts and claim decisions
Source snapshot Request URLs, query parameters, source hashes and normalised public fields
Validated profiles Five real Steam profiles conforming to the draft schema
Discovery queries Ten deterministic hard and soft constraint tasks
Adversarial fixtures Four source, freshness, duplication and prompt-injection tests
Raw results Clean and adversarial rankings with summary measures

The build and evaluation scripts are build_public_case_study.py and run_public_case_study.py. Their dependencies are pinned in public-case-study-requirements.txt.

The schema does not yet include:

  • a UCP extension definition;
  • an ACP feed mapping;
  • an RDF serialisation;
  • SHACL shapes;
  • a canonical game-identity registry;
  • a controlled vocabulary for claim properties and units;
  • conformance tests across independent implementations.

Those are implementation tasks beyond the current case study.

6. Private telemetry and ranking stay behind controlled interfaces

6.1 Raw telemetry is not a discovery feed

Session logs can help estimate:

  • session duration;
  • interruption points;
  • party-size distribution;
  • return intervals;
  • mode population.

Publishing raw events would expose player behaviour and create a broad data surface. The marketplace should compute bounded aggregates and release only the fields needed for the claim.

The draft example uses a cohort threshold and user-level contribution bounds. It explicitly says that no formal differential privacy claim is made.

Calling a system privacy-preserving requires more. NIST guidance expects a defined privacy unit, adjacency relation, privacy parameters and composition accounting.[43] Small cohorts and repeated releases need particular care.

Privacy controls can also reduce coverage for niche games. Suppression rules may remove the very long-tail evidence an agent needs. Evaluation should measure this effect by popularity cohort.

6.2 Entitlement checks should return the minimum useful answer

A group query can use marketplace-issued tokens scoped to:

candidate set
authorised participants
required edition
expiry time

Each marketplace evaluates its own accounts and returns an aggregate result. The agent learns whether the group constraint passed. It does not receive the underlying libraries.

Cross-marketplace ownership remains difficult because editions and bundles do not share one canonical identifier. The profile can carry mappings, but those mappings need governance and correction.

6.3 Ranking weights need not be public

Publishing exact recommendation weights would make manipulation easier and could reveal commercially sensitive systems.

The profile asks for a smaller disclosure:

  • who generated the candidate;
  • whether placement was organic or commercially influenced;
  • which declared constraints were satisfied;
  • which evidence records support the explanation;
  • a trace identifier for audit.

The platform can retain private features, fraud models and experiment assignment.

6.4 Agents should disclose their own transformation

An agent may combine:

  • marketplace candidates;
  • user preferences;
  • external reviews;
  • third-party duration estimates;
  • its own inferred attributes.

Its answer should distinguish marketplace-provided claims from agent-added evidence. The agent should also identify the model and method used to extract new claims.

Without that boundary, a marketplace can be blamed for an inference it never supplied.

7. Threat model

7.1 Publisher overstatement

Once agents read structured discovery fields, publishers gain an incentive to optimise those fields.

A developer can label a game:

relaxing
deep
easy to return to
ideal for short sessions

The labels may be sincere and still be too broad. They may also be written to trigger agent filters.

The control is source separation. A developer declaration remains visible as a developer declaration. Platform verification and population evidence use different claim classes.

7.2 Review and profile attacks

Shilling attacks inject profiles or interactions to push or demote items. A recent survey covers decades of poisoning methods and countermeasures. An agent-facing evidence layer does not remove this attack surface.[38]

Game-marketplace variants include:

  • coordinated positive or negative review bursts;
  • plausible accounts with fabricated play histories;
  • helpfulness-vote manipulation;
  • duplicated claims presented as independent sources;
  • targeted attacks against long-tail games.

The profile helps an evaluator retain source and time windows. Detection still requires platform abuse systems and adaptive attack testing.

7.3 Ranking manipulation in retrieved text

Research on conversational search shows that adversarial product-page text can raise the rank of a target item. The attack does not require control over the ranking model.[39]

An agent that reads store pages, patch notes and reviews also faces indirect prompt injection. InjecAgent demonstrates this risk for tool-integrated agents.[40]

Retrieved text must be treated as data. It should never change tool authority or system instructions.

7.4 Record forgery and unqualified trust in the record

A profile record asserts its own provenance. A relay can emit marketplace_verification evidence without performing a test. Free-text fields such as method, cohort and uncertainty.description also reach the agent as readable content.

Source classes are only as trustworthy as the channel that delivered them. A deployment needs authenticated transport or signed records bound to the issuer. The agent must treat every string in a record as data, never instruction.

An issuer can also omit unfavourable evidence. Status and supersession history help with correction, while completeness still requires comparison against the catalogue or another accountable source.

7.5 Stale evidence

A claim can remain accurate for the source snapshot and wrong for the current release.

Attackers may exploit this by:

  • quoting favourable pre-patch reviews;
  • retaining old compatibility claims;
  • presenting a removed mode as current;
  • applying one target platform's evidence to another edition.

Release identity and expiry rules reduce the error. They do not decide the correct lifespan automatically.

7.6 Privacy inference

Even aggregate responses can leak information when:

  • the group is small;
  • queries are repeated with one member changed;
  • the agent controls candidate sets;
  • rare ownership patterns identify a person.

Entitlement endpoints need query budgets, minimum group sizes and protections against differencing attacks. A boolean response is not automatically private.

7.7 Hidden commercial influence

An affiliate relationship can affect candidate selection or ordering. Undisclosed influence creates a conflict between the user's request and the agent's incentives.

The profile can mark known placement status. Enforcement depends on platform and regulatory policy.

7.8 Threat and control summary

Threat Profile control Remaining platform work
Publisher overstatement Source-qualified claim class Verification and sanctions
Review poisoning Time windows and source records Detection and moderation
Prompt injection Typed evidence references Agent isolation and tool policy
Record forgery Issuer and claim class Authenticated transport or record signing
Selective omission Status and supersession history Completeness audit against the catalogue
Stale evidence Release scope and expiry Revalidation triggers
Ownership leakage Aggregate result shape Query budgets and privacy controls
Hidden sponsorship Placement field Audit and commercial enforcement
Cross-edition mismatch Work, edition and release identity Canonical mapping governance

8. Public Steam case study

8.1 Scope

The case study tests one bounded question:

Does source, release and evidence validation change deterministic discovery behaviour when the inputs are incomplete or adversarial?

It does not reproduce a Steam recommender. It uses public data for five games:

Game AppID Selected role in the case study
Portal 2 620 Local and online co-op with player-report evidence
Deep Rock Galactic 548430 Online co-op with short-session, co-op and return evidence
Hades 1145360 Single-player short-session evidence
Balatro 2379780 Single-player short-session evidence
Left 4 Dead 2 550 Online co-op with player-report and return evidence

The selection is purposeful. It creates different combinations of operating system support, controller support and multiplayer modes. It is not intended to represent the Steam catalogue.

8.2 Public source capture

The build script retrieves:

  • public store metadata through Valve's appdetails endpoint;
  • current connected-player counts through the documented Steam Web API;
  • reviews through the documented review endpoint.

The appdetails endpoint is Valve-hosted and undocumented. The snapshot records that status, and the generated profiles label it as an unstable extraction transport. It is not treated as a normative API.

The review query fixes:

language = english
purchase_type = steam
filter = all
review_type = all
day_range = 365
num_per_page = 100
filter_offtopic_activity = 1

The source snapshot records the review endpoint, fixed query parameters, per-page cursor hashes and response hashes. It stores 15 short, human-checked excerpts and their review fingerprints. Steam account identifiers and recommendation IDs are discarded. The excerpts and fingerprints may still be linked to the underlying public reviews, so the records are not anonymised.

8.3 Real profiles with explicit limits

The build generated five records that validate against the draft schema. Each record contains:

  • target-platform support;
  • full-controller support;
  • single-player, online co-op and local co-op flags;
  • a point-in-time connected-player count;
  • three experiential claim slots;
  • a current US offer snapshot.

The experiential claims cover short-session fit, reported co-op fit and returning after time away. Each known claim links to one or more selected review excerpts. An unknown or not-applicable result remains explicit.

The records use a public snapshot identifier because no public build identity was available through the selected sources. The identifier binds claims to the case-study acquisition, not to a verified executable build.

External-launcher status remained unknown for every offer. The source set did not provide a dependable field, so the profiles did not infer one.

8.4 Deterministic discovery comparison

The query set contains ten tasks. Hard constraints cover target platforms, controller support and multiplayer modes. Soft constraints use the three experiential properties.

The comparison uses two deterministic engines:

Engine Behaviour
Flattened baseline Uses matching values without checking claim class, evidence source type or release scope, and counts every matching soft claim
Profile-aware Requires a functional-fact claim class, an approved evidence source type and current release scope for hard facts, preserves unknowns and deduplicates review evidence

Both engines use the point-in-time connected-player count only as a final tie-break. It is not treated as evidence of quality.

The two engines agreed on all ten clean queries and returned the same top result on the nine that produced a candidate. The clean tasks therefore do not establish an improvement from the profile.

One query required a known absence of an external launcher. Both engines returned no candidate because all five offers carried an explicit unknown. That outcome demonstrates refusal to invent a hard fact.

8.5 Adversarial fixtures

The fixture set applies four mutations:

  1. a model-inferred claim says Balatro supports online co-op;
  2. a stale Hades claim says an unrelated release supports online co-op;
  3. one Portal 2 review is copied into three short-session claims;
  4. a Deep Rock Galactic review excerpt contains an instruction to ignore the ranking policy.

Each mutated record still validates structurally. The test therefore examines source policy and release checks under valid JSON.

8.6 Measured results

The raw results report:

Measure Flattened baseline Profile-aware engine
Target promotions across four fixtures 3 0
Top recommendations violating a hard constraint 1 0
Rank change from the prompt-injection string 0 0

The forged Balatro claim moved an ineligible game from absent to first place in the flattened baseline. The profile-aware engine rejected the claim because its claim class was a model inference rather than a functional fact, and because model inference is not an approved evidence source type for a hard fact.

The stale Hades claim entered the flattened candidate list at rank two. The profile-aware engine rejected its release scope.

Duplicated Portal 2 evidence moved the game from rank three to rank one in the flattened baseline. The profile-aware engine aggregated unique review evidence keys across matching claims and left the rank unchanged.

A counterfactual control replaced the repeated key with three distinct synthetic keys. Portal 2 then moved from rank three to rank two in the profile-aware engine. This isolates evidence-key deduplication from the experience-evidence threshold.

The prompt-injection string changed neither deterministic engine. This result does not establish LLM-agent resistance. It confirms only that the case-study runner never interprets evidence text as an instruction.

8.7 Reproduction and limits

The bundle can be verified independently with:

python -c "import hashlib;print(hashlib.sha256(open('game-discovery-evidence-bundle.zip','rb').read()).hexdigest())"

After extracting the bundle, the case-study artifacts can be regenerated with:

pip install -r public-case-study-requirements.txt
python build_public_case_study.py
python run_public_case_study.py

The source APIs are live. A later run can produce different prices, player counts, review pages or edited review text. The build also stops with an error if a selected review is edited, removed or falls outside the fixed 365-day window, so the annotation set has to be refreshed before a later rebuild. The captured snapshot and its hashes are the evidence for the reported result.

This case study does not measure player satisfaction. It does not compare learned recommenders or evaluate private telemetry. Five games cannot establish catalogue-wide value.

Its result is narrower: the proposed source, release and evidence-deduplication rules changed the behaviour of a deterministic discovery pipeline under three data-manipulation fixtures, while explicit unknown handling left one unsupported query unanswered in both engines.

9. Proposed wider evaluation

9.1 Task set

The benchmark should include:

  • factual compatibility requests;
  • group ownership and entitlement constraints;
  • short-session requests;
  • co-op coordination constraints;
  • return-after-break requests;
  • long-tail alternatives;
  • deliberately underspecified requests;
  • requests that no candidate can satisfy.

Experiential tasks need human-reviewed definitions. Low commitment should be decomposed into session length, onboarding burden and progression pressure.

9.2 Baselines

Compare:

Baseline Available evidence
Popularity Review count, current players or another declared popularity signal
Store metadata Factual metadata and tags
Review retrieval Metadata plus retrieved review passages
Agent without profile Multi-source model with no source-qualified schema
Profile without freshness Source-qualified claims without expiry
Full profile Source, scope, freshness and uncertainty

Attack handling is applied identically to every arm. This keeps the defence constant and isolates the profile.

Existing Steam recommender research provides useful accuracy and diversity baselines. Cheuque and colleagues compared several recommendation models on Steam data. CPGRec explicitly addresses the balance between accuracy and long-tail exposure.[31][32]

The benchmark should use tuned simple baselines. A weak popularity baseline would make a complex system look better without establishing practical value.

9.3 Ground truth

Reviews used to derive a claim cannot also be the sole ground truth for that claim.

Simulated aggregate records cannot be generated from the ground-truth labels for the same task. Tasks that depend on simulated telemetry should be reported separately from the headline comparison.

Use:

  • official or marketplace verification for functional facts;
  • held-out human annotation for extracted claims;
  • defined player instruments for population experience;
  • post-play participant judgement for personalised fit;
  • temporal snapshots for freshness.

PXI and challenge-specific instruments can support selected constructs. They should not be combined into an unvalidated universal score.[28][29]

9.4 Adversarial suite

The test set should include:

  • exaggerated developer fields;
  • coordinated review bursts;
  • duplicated evidence;
  • stale pre-patch evidence;
  • cross-target-platform evidence substitution;
  • prompt injection inside a review;
  • hidden affiliate placement;
  • privacy differencing against entitlement checks.

Attackers should know the validation rules. Testing only malformed records does not establish resistance to strategic input.

9.5 Measures

Headline measures:

Measure Definition
Hard-constraint violation Recommended items failing a required factual or eligibility constraint
Unknown-field honesty Missing facts returned as unknown rather than inferred
Claim support Output claims linked to evidence that entails them
Provenance completeness Required source, method and scope fields present
Freshness error Claims applied outside their release or validity window
Calibration Agreement between stated probability and observed outcome
Manipulation lift Target rank change under a defined attack budget
Long-tail exposure Exposure by item-popularity cohort
Decision fidelity Rank change when the stated reason is removed
Privacy disclosure Information released beyond the authorised aggregate
Production cost Records maintained per title, plus latency and token overhead against the no-profile arm
Outcome stability Variation in the other headline measures across repeated runs of the same task

Popularity-bias research warns that more long-tail exposure is not automatically better for users.[41] This evaluation therefore reports long-tail exposure alongside hard-constraint violation and the post-play outcomes defined in section 9.6.

Each task should be run repeatedly. The headline comparison should report a distribution across runs.

9.6 Human study

A later study can compare:

  • conventional recommendations;
  • recommendations with evidence citations;
  • recommendations with source disagreement and uncertainty.

Participants should make a real choice and play the selected game. Outcomes should include constraint satisfaction, verification time and override behaviour. Self-reported trust alone is insufficient.

9.7 Reproducibility

The repository should publish:

  • schema and example records;
  • source acquisition manifest;
  • source snapshots where licences allow;
  • identity mappings and correction history;
  • annotation guide and adjudication record;
  • task set and split definition;
  • attack fixtures and budgets;
  • model and prompt manifests;
  • append-only decision traces;
  • evaluation scripts;
  • raw and summarised results.

FEVR provides a useful framework for recommender-system evaluation and reporting.[44] This study should preregister confirmatory outcomes and separate them from exploratory analysis.

10. The strongest counterargument

The binding objection is commercial. Amazon blocked Muse instead of publishing a richer interface to it. A marketplace that exposes source-qualified evidence lowers the cost of comparing its catalogue against a competitor's and receives ranking scrutiny in return.

A second objection is about cost: an agent can already read store pages, reviews and public APIs. A new profile adds implementation cost and risks turning subjective play into rigid categories.

The case study supports part of that objection. The flattened baseline and the profile-aware engine returned the same top result on every clean query that produced a candidate. The profile changed behaviour only when evidence was incomplete or manipulated.

Many low-stakes queries do not need a profile. An agent can summarise public pages and let the player decide. Marketplace recommendations may already outperform an external agent because the platform holds stronger behavioural data.

The profile is justified where one or more of these conditions hold:

  • the request contains a hard compatibility or entitlement constraint;
  • the explanation needs auditable source boundaries;
  • evidence changes by release, region or cohort;
  • several marketplaces need comparable records;
  • manipulation risk makes unqualified text unsafe.

The profile should not standardise taste. It should describe the evidence used to make a contextual prediction.

An implementation that forces every game into one fixed experience taxonomy would fail. The vocabulary needs versioning, extension points and an explicit unknown state.

The marketplace may also decide that the value does not justify exposing an external interface. A game-domain profile can still improve the marketplace's own agent and internal discovery services.

11. What this paper does not claim

This paper does not claim:

  • that agent-mediated commerce replaces conventional storefronts;
  • that external agents should receive unrestricted access to a marketplace;
  • that UCP, ACP or AP2 lack product-discovery support;
  • that the proposed profile is the first provenance or agent-recommendation architecture;
  • that the term Evidence Contract originates here;
  • that public Steamworks documentation describes Valve's private ranking systems;
  • that Steam uses one recommendation algorithm across every discovery surface;
  • that tags, reviews or playtime are objective measures of game quality;
  • that one experience taxonomy can describe every player;
  • that source provenance proves a claim is true;
  • that a citation proves a ranking explanation is faithful;
  • that thresholded telemetry is differentially private;
  • that the schema prevents review poisoning or prompt injection;
  • that the profile authenticates the party asserting a claim;
  • that long-tail exposure always improves player outcomes;
  • that cross-store game identity or entitlements have been solved;
  • that selected review excerpts estimate the wider player population;
  • that the five-game case study establishes recommendation quality;
  • that its prompt-injection fixture establishes LLM-agent resistance;
  • that the proposed wider benchmark or player study has been implemented.

The paper proposes a profile, publishes an initial schema and reports a small public-data case study. It also defines the wider work needed to evaluate player outcomes.

12. Conclusion

Agentic commerce already has catalogue, checkout and payment protocols. The game-marketplace problem begins where those generic layers stop.

A useful game recommendation can depend on edition identity, private entitlements and release-specific compatibility. It can also depend on population evidence about sessions or cooperative play. These inputs have different owners and different uncertainty.

Steam's public systems make that diversity visible. Tags combine input from several contributor classes, while Deck compatibility rests on a marketplace test against published criteria. Reviews retain purchase and playtime context under a separate score policy. Valve's published 2019 and 2020 descriptions said the Interactive Recommender used behavioural patterns and stayed separate from tag-driven discovery.

An external agent should not flatten those records into one confident description.

The proposed Game Discovery Profile binds claims to a game edition and release. It separates functional facts, marketplace verification, aggregate observations, population experiences, model inferences and personalised predictions into distinct claim classes, and records the source behind each one. It records applicability and method, and supports private eligibility checks with aggregate answers. It marks commercial placement while leaving ranking weights private.

The profile is a synthesis over existing standards and research. Its value has to be established through implementation and comparative evaluation.

The public case study establishes one narrower result. Source, release and evidence-deduplication checks rejected three adversarial target promotions that a flattened baseline accepted. Explicit unknown handling also left a launcher constraint unanswered by both engines. The sample does not establish better recommendations for players.

The wider test is practical: a player states a hard constraint, and the returned candidates either satisfy it or say that the evidence is missing.

The storefront keeps the marketplace functions, and the agent becomes another interface to its evidence.

About the author

Jason Doyle writes about software systems, reliable products, applied AI and practical controls for systems that influence human and organisational decisions. He publishes at jasondoyle.ie and can be contacted at [email protected].

References

  1. Meta, Introducing Muse: The World's First Personal AI Agent Built for Everyone, 8 September 2026, https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/.

  2. Todd Bishop, Amazon blocks Meta's Muse AI assistant in new standoff over agentic shopping, GeekWire, 20 September 2026, https://www.geekwire.com/2026/amazon-blocks-metas-muse-ai-assistant-in-new-standoff-over-agentic-shopping/.

  3. Shopify, Meta is now an AI channel in your admin, 8 September 2026, and Selling on Meta, accessed 28 September 2026, https://changelog.shopify.com/posts/meta-is-now-an-ai-channel-in-your-admin, https://help.shopify.com/en/manual/online-sales-channels/agentic-storefronts/meta.

  4. Elizabeth Howcroft, Reuters, Banks warn AI shopping bots raise scam, fraud and data-privacy risks, 22 September 2026, https://www.reuters.com/legal/litigation/banks-warn-ai-shopping-bots-raise-scam-fraud-data-privacy-risks-2026-09-22/.

  5. ASB Bank Ltd., Bank of America, Capital One, Commonwealth Bank of Australia, ING Group and NatWest Group, Building Trust in Agentic Commerce, 22 September 2026, https://www.commbank.com.au/content/dam/commbank-assets/business/latest/2026/building-trust-in-agentic-commerce-industry-report-final.pdf.

  6. Universal Commerce Protocol, Official Specification, release 25 August 2026, https://ucp.dev/2026-08-25/specification/overview/.

  7. OpenAI and Stripe, Agentic Commerce Protocol, beta project, stable specification snapshot 17 April 2026, https://github.com/agentic-commerce-protocol/agentic-commerce-protocol/tree/main/spec/2026-04-17.

  8. Google Agentic Commerce, Agentic Payment Protocol v0.2, 28 April 2026, https://github.com/google-agentic-commerce/AP2/blob/v0.2.0/docs/ap2/specification.md.

  9. Visa, Specifications - Trusted Agent Protocol, accessed 28 September 2026, https://developer.visa.com/capabilities/trusted-agent-protocol/trusted-agent-protocol-specifications.

  10. Visa, Visa to enable AI-driven shopping for businesses, 8 April 2026, https://corporate.visa.com/en/sites/visa-perspectives/newsroom/visa-intelligent-commerce-connect-ai-shopping-for-businesses.html.

  11. Mastercard, Mastercard gives merchants a simpler way to build, connect and scale AI-powered shopping experiences, 9 September 2026, https://www.mastercard.com/us/en/news-and-trends/press/2026/september/mastercard-gives-merchants-a-simpler-way-to-build--connect-and-s.html.

  12. Valve, Steam Tags, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/store/tags.

  13. Valve, Update to Store Tags: Additions, Removals, and Edits, 18 May 2026, https://steamcommunity.com/ogg/593110/announcements/detail/673994309884707519.

  14. Valve, Visibility on Steam, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/marketing/visibility.

  15. Valve, Introducing the Interactive Recommender, 11 July 2019, https://steamcommunity.com/ogg/593110/announcements/detail/1612767708821405787.

  16. Valve, Introducing The Steam Interactive Recommender, 18 March 2020, https://steamcommunity.com/ogg/593110/announcements/detail/1716373422378712841.

  17. Valve, Personalized Shopping With New Tag, Genre, and Category Pages, 7 September 2022, https://steamcommunity.com/ogg/593110/announcements/detail/3091162528094367314.

  18. Valve, User Reviews, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/store/reviews.

  19. Valve, User Reviews - Get List, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/store/getreviews.

  20. Valve, User Reviews Revisited, 15 March 2019, https://steamcommunity.com/ogg/593110/announcements/detail/1808664240333155775.

  21. Valve, Update to User Reviews: New Helpfulness System, 14 August 2024, https://steamcommunity.com/ogg/593110/announcements/detail/4326355263805583416.

  22. Valve, New Profile Privacy Settings, 10 April 2018, https://steamcommunity.com/games/593110/announcements/detail/1667896941884942467.

  23. Valve, Steam Deck and Steam Machine Compatibility Review, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/steamhardware/compat.

  24. Valve, Accessibility Features, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/accessibility_features.

  25. Schema.org, VideoGame, SoftwareApplication, Product, Offer and Review, accessed 28 September 2026, https://schema.org/VideoGame, https://schema.org/SoftwareApplication, https://schema.org/Product, https://schema.org/Offer, https://schema.org/Review.

  26. Jin Ha Lee et al., Developing a Video Game Metadata Schema for the Seattle Interactive Media Museum, International Journal on Digital Libraries, 2013, https://doi.org/10.1007/s00799-013-0103-x.

  27. Jose P. Zagal et al., Towards an Ontological Language for Game Analysis, Digital Games Research Association Conference, 2005, https://doi.org/10.26503/dl.v2005i1.136.

  28. Vero Vanden Abeele et al., Development and validation of the Player Experience Inventory: A scale to measure player experiences at the level of functional and psychosocial consequences, International Journal of Human-Computer Studies, 2020, https://doi.org/10.1016/j.ijhcs.2019.102370.

  29. Alena Denisova et al., Measuring perceived challenge in digital games: Development and validation of the challenge originating from recent gameplay interaction scale (CORGIS), International Journal of Human-Computer Studies, 2020, https://doi.org/10.1016/j.ijhcs.2019.102383.

  30. Dayi Lin et al., An empirical study of game reviews on the Steam platform, Empirical Software Engineering, 2019, https://doi.org/10.1007/s10664-018-9627-4.

  31. German Cheuque, Jose Guzman and Denis Parra, Recommender Systems for Online Video Game Platforms: the Case of STEAM, The Web Conference Companion, 2019, https://doi.org/10.1145/3308560.3316457.

  32. Xiping Li et al., Category-based and Popularity-guided Video Game Recommendation: A Balance-oriented Framework, The Web Conference, 2024, https://doi.org/10.1145/3589334.3645573.

  33. Haohan Yuan et al., Position: Recommender Systems Should Move Beyond Platform-Centric Ranking toward Personal Agent-Mediated Recommendation, arXiv preprint, 21 July 2026, https://arxiv.org/abs/2609.11942.

  34. Ruoyu Xu, Gaoxiang Li and Victor S. Sheng, GAVEL: Evidence-Contract Debate with Mechanized Scrutiny for Provenance-Grounded Fact-Checking, Findings of ACL 2026, https://doi.org/10.18653/v1/2026.findings-acl.1789.

  35. W3C, PROV-O: The PROV Ontology, W3C Recommendation, 30 April 2013, https://www.w3.org/TR/prov-o/.

  36. W3C, Web Annotation Data Model, W3C Recommendation, 23 February 2017, https://www.w3.org/TR/annotation-model/.

  37. W3C, Shapes Constraint Language (SHACL), W3C Recommendation, 20 July 2017, https://www.w3.org/TR/shacl/.

  38. Thanh Toan Nguyen et al., Manipulating Recommender Systems: A Survey of Poisoning Attacks and Countermeasures, ACM Computing Surveys, 2024, https://doi.org/10.1145/3677328.

  39. Samuel Pfrommer et al., Ranking Manipulation for Conversational Search Engines, EMNLP 2024, https://doi.org/10.18653/v1/2024.emnlp-main.534.

  40. Qiusi Zhan et al., InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents, Findings of ACL 2024, https://doi.org/10.18653/v1/2024.findings-acl.624.

  41. Anastasiia Klimashevskaia et al., A survey on popularity bias in recommender systems, User Modeling and User-Adapted Interaction, 2024, https://doi.org/10.1007/s11257-024-09406-0.

  42. Ingrid Nunes and Dietmar Jannach, A systematic review and taxonomy of explanations in decision support and recommender systems, User Modeling and User-Adapted Interaction, 2017, https://doi.org/10.1007/s11257-017-9195-0.

  43. NIST, Guidelines for Evaluating Differential Privacy Guarantees, Special Publication 800-226, 2025, https://doi.org/10.6028/NIST.SP.800-226.

  44. Eva Zangerle and Christine Bauer, Evaluating Recommender Systems: Survey and Framework, ACM Computing Surveys, 2022, https://doi.org/10.1145/3556536.

  45. Amazon, Conditions of Use, updated 14 August 2026, https://www.amazon.com/gp/help/customer/display.html?nodeId=GLSBYFE9MGKKQXXM.

  46. University of Washington GAMER Group, Video Game Metadata Schema v4.2, 12 December 2024, https://github.com/uwgamergroup/video-game-metadata-schema/blob/main/VGMS_v4.2_20241212.pdf.

  47. Valve, Building and Editing Store Pages, Steamworks Documentation, accessed 28 September 2026, https://partner.steamgames.com/doc/store/page.

  48. Valve, ISteamUserStats: GetNumberOfCurrentPlayers, Steamworks Web API, accessed 28 September 2026, https://partner.steamgames.com/doc/webapi/ISteamUserStats#GetNumberOfCurrentPlayers.

  49. Yaxin Zhu et al., Faithfully Explainable Recommendation via Neural Logic Reasoning, NAACL 2021, https://doi.org/10.18653/v1/2021.naacl-main.245.