# Ansmeter — full text corpus > Complete text dump of the Ansmeter knowledge base and reading list, for AI ingestion. Ansmeter is applied research into how AI systems (ChatGPT, Gemini, Perplexity, Claude) see, interpret, and recommend brands. The link-first index is at https://ansmeter.com/llms.txt. --- Source: https://ansmeter.com/knowledge-base/ansmeter-corpus-guide # Guide to the Ansmeter Knowledge Base ## What you are looking at This is not a blog or a news feed. This is a **research library** — a structured corpus of texts about how brands exist in the responses of AI systems. Or don't exist. Each text occupies a specific place: some introduce concepts, others examine specific phenomena, and others provide practical tools. All materials are cross-linked and organized into reading paths — ready-made sequences for a specific role or task. The thematic field of the corpus: how a model represents a brand internally, where it gets its information, how it forms recommendations, why the same brand looks different across systems and languages, and what to do about it. --- ## How each article is structured All materials in the corpus follow a uniform structure. This is intentional — consistency makes it easy to navigate and compare texts. ### Attribute line Above the title — a compact line: material type, ●● difficulty level, reading time, and the tasks the article helps solve. More on types and levels below. ### Description and lead quote Below the title — one or two sentences explaining the essence. Below that:
An italicized lead quote that sets the tone and formulates the central thesis of the article.
### Research card Most articles include a table with three fields:
Research question
The specific question the text answers
Evidence type
What supports the conclusions: academic papers, platform data, Ansmeter's own test runs, or a combination
Data freshness
The period to which the facts cited in the article apply
### Body text and table of contents The text is divided into sections with subheadings. On the right — a table of contents for quick navigation between sections. ### Three closing blocks Research articles end with three blocks that capture the current state of knowledge on the topic:
What seems well established

Conclusions supported by reproducible data and confirmed from multiple sources.

What remains uncertain

Questions without a clear answer, platform dependencies, immature metrics.

What this changes in practice

Concrete implications for a brand: what to do, what to change, what to watch.

### Sources and related materials At the end — a numbered list of sources. Numbers in square brackets [1], [2] appear throughout the text and link to the corresponding entry. After the sources — a block with links to other texts in the corpus on the same topic. --- ## Material types Each text is marked with a colored badge in the attribute line. Foundational text — the foundation of the corpus. Introduces key concepts, explains mechanisms, builds the groundwork. If you're just starting — you'll most likely begin with foundational texts. Research article — an analytical article backed by data. Examines a specific phenomenon, always contains a research card and three closing blocks. Field note — an observation recorded in real time. Less formalized than a research article; captures a discovered effect or anomaly with primary data. Update — an analysis of a specific platform change. Tied to a date, includes an assessment of how the change affects brand visibility. Reference — material for constant reference. A reference of all terms and metrics, convenient to return to while reading any other article. Observation template — a practical tool. A card that can be used to record observations for each study. Guide — a navigational text (including this one). Helps orient within the corpus and choose an entry point. --- ## Difficulty level Each material is marked with colored dots — in the material list and in the attribute line on the article page. Introductory — an entry point. No prior knowledge required. Suitable for a first encounter with the topic. ●● Intermediate — assumes you've already read at least a couple of introductory texts and understand the corpus's core concepts. ●●● Advanced — for readers who have completed at least one reading path. Touches on methodological nuances, edge cases, and non-obvious relationships. --- ## Reading paths A path is a ready-made sequence of texts selected for a role or task. You don't have to complete a path in full: even the first two or three texts provide a working understanding of the topic. **Getting started** — the minimum set. What AI visibility is, where the model gets information about a brand, and what the customer journey through an AI intermediary looks like. **For the business owner** — focus on business decisions: the economics of invisibility, the competitive landscape, when and why to run a study. **For the marketer** — sources, citations, what transfers from SEO and what doesn't, practical steps to grow visibility. **For the technical lead** — infrastructure: machine-readable data, markup, access control, integrations. **For the researcher** — methodology: how the benchmark works, which metrics and why, limitations, reproducibility. **For the agency and consultant** — how to explain the topic to a client, which arguments work, how to structure diagnostics. **Full course** — all materials in the corpus in recommended order, from introduction to advanced topics. --- ## Navigating the knowledge base The knowledge base page is organized into four modes. Switch between them using the cards at the top of the page. ### All materials The full catalog in table form. Each row shows the article title with a brief description, type badge, colored difficulty dots, and reading time. The table can be sorted by any column (click the header) and filtered through the search bar — search works across titles and descriptions. ### Find by task Filtering mode. At the top — task cards: "Understand the problem", "Start diagnostics", "Assess risks", "Prepare implementation", and others. Select a task — get only the materials that address it. Additional filters let you narrow the selection by topic and difficulty level. ### Reading path Seven ready-made paths — sequences of texts selected for a role or task. Each path expands on click and shows numbered steps with types, levels, and reading times. Total time and expected outcome are indicated. The "Start reading" button opens the first text in the path. ### Reference Three reference blocks convenient to return to while reading: - **Concepts** — reference of all corpus terms. Each term includes a definition and an italicized practical meaning — what this term means for decision-making. - **Report metrics** — weight tables for the main score and diagnostic metrics. Show what makes up the AI Visibility Score and how to interpret each indicator in the report. - **Research scenarios** — types of questions Ansmeter asks the model during testing. Explain what each scenario tests and how it affects the final score. --- *The Ansmeter corpus is available in five languages: Russian, English, Spanish, French, and German.* --- Source: https://ansmeter.com/knowledge-base/mini-research-card # Mini-Research Card for the Ansmeter Database Below is the minimum observation standard from which one can begin building a cumulative database across brands, categories, and answer systems. The card is intentionally simple: it should be reproducible and suitable both for desk research and for regular monitoring within a product or marketing team. The card's core principle is to record not only the fact that a brand appears, but the entire contour of the answer: the wording of the question, the user's intent, the language, the system, the type of sources, the brand's role, the presence of citations, the nature of category drift, the update lag, and the final practical interpretation. Otherwise, the database will quickly turn into a repository of attractive screenshots with no analytical depth. ## Card fields ### System and mode For example: Google AI Overviews, AI Mode, ChatGPT Search, Copilot Search, Perplexity. ### Date and geography of the research run Record the date, locale, interface language, and country if it affects the results. ### Original question The exact wording of the user's question, without editorial changes. ### Intent type Informational, comparative, commercial, local, navigational, research-oriented. ### Brand's role in the answer Absent; mentioned; cited; shapes the frame of the answer; appears in the shortlist. ### Source type and quality Official, editorial, user-generated, institutional, catalog; strong or weak. ### Signs of category drift Whether there is task drift, a shift in the language of comparison, or the insertion of unrelated alternatives. ### Signs of update lag Whether the answer matches the latest known facts and where the delay likely came from. ### Editorial conclusion A brief interpretation: where the problem lies, where the strong area is, and what to check next. --- Source: https://ansmeter.com/knowledge-base/why-strong-brands-become-invisible # Article 1. Why a Strong Brand Can Be Invisible to Answer Systems *Overview and methodology article.* **Research question.** Why a brand that is recognizable — and even beloved by people — may turn out to have low machine distinctness for an answer system at the moment of real choice. **Evidence type.** Research on language-model interpretability, documents from search and answer platforms, and market data on the scale of AI search. **Data freshness.** Current product and platform data are presented as of March 2026. > Brand recognition and machine distinctness are not the same thing. For an answer system, what matters is not simple name recognition, but the stability of the whole entity: the name, the category, the properties, the relationships, and the external validation. ## From recognition to machine distinctness The most deceptive mistake in discussions of digital visibility today is that many people still think in the categories of classic search. If a brand is well known, if people search for it by name, if it has a strong site, meaningful direct traffic, and a stable media reputation, it seems natural to assume that it will be just as obvious to modern answer systems. But this is exactly where the new environment starts to resist the old logic. For a human, a strong brand is a name, a reputation, a web of associations, and accumulated trust. For an answer system, that is not enough. It needs not merely to “know that the brand exists,” but to be able to confidently assemble it in an answer: distinguish it from similar entities, assign it to the right category, connect it to relevant properties, verify it against external sources, and restate it without semantic distortion. That is what creates the paradox that is already becoming a business problem. A brand can be highly visible to people and at the same time poorly distinguishable to machines. It can have strong demand across the conventional web, yet remain a secondary player in answers from ChatGPT, Google AI Overviews, Gemini, Copilot, or Perplexity. And the issue is not some single technical defect. More often, the cause runs deeper: human recognition and machine readability are not the same thing. Modern large language models do not store knowledge about a company as a neat card. Research from recent years shows that factual associations are distributed across model parameters, intermediate computations, and, in many systems, the external documents that are mixed in at answer time [1][2][3]. That means the brand exists inside the machine not as a coherent object, but as a pattern of associations: the name, adjacent terms, the category, typical properties, competitors, usage scenarios, fragments of reputation, traces of citation, and probabilistic expectations about what usually follows its name. That construction can be strong, or it can be fragile. And fragility matters here most of all. ## Four reasons for machine invisibility It emerges for several reasons. The first is entity ambiguity. If a company uses several names, describes the product differently on different pages, mixes its corporate and consumer name, or operates in a category where the same term has many meanings, the model receives not a stable entity but a set of partially overlapping signals. A human will usually disentangle that ambiguity on their own. A machine does it worse, especially when it has to answer quickly and briefly. The second reason is the gap between self-description and external validation. It is natural for a brand to describe itself in favorable terms: “leading platform,” “innovative service,” “solutions ecosystem.” But in answer scenarios, answer systems increasingly rely not only on their own knowledge, but also on external web sources. Google states explicitly that its AI features fan a query out into subtopics and multiple data sources, and then select supporting links [4]. OpenAI describes ChatGPT Search as a mechanism for producing current answers grounded in web sources [5]. Perplexity puts it even more simply: the system searches the internet in real time and then compresses what it finds into a short answer [6]. For a brand, this implies an unpleasant but important fact: its own site is no longer the sovereign source of truth about itself. It is only one voice in a much broader chorus. The third reason is semantic diffusion. Many strong companies are broadly present across the internet, but not in a coordinated way. One part of the material is written in the language of sales, another in the language of technical documentation, a third in the language of press releases, and a fourth in the language of customer reviews. For a human, that is a natural polyphony. For AI, it often means an unstable center of gravity. The model may remember the brand name, but connect it only weakly to a specific task. It may place the company in the right industry, yet fail to understand how it differs from competitors. It may reproduce an old positioning statement and miss the new one. And sometimes it simply “stitches together” the brand from fragments drawn from different sources, with the result that the defining properties are not the ones the company itself considers most important. The fourth reason is the limited and fragile nature of machine memory itself. Survey research on the mechanics of knowledge in language models emphasizes that parametric knowledge in such systems is distributed, prone to staleness, and sensitive to the wording of the question [3][7]. In other words, the model may “know” about a company, yet fail to retrieve that knowledge in the right phrasing. Or retrieve it only in fragments. Or blend it with a neighboring entity. For a consumer-facing answer, this is especially dangerous: the user does not see the model’s internal uncertainty, but only the finished summary. The error appears not as hesitation, but as a confident and inaccurate interpretation. ## Functional visibility matters more than the name alone That is why a strong brand often becomes machine-invisible not in an absolute sense, but in a functional one. It may be known by name, yet not recommended when the user asks about a class of solutions. It may be mentioned, but without its key advantages. It may be cited, but only on secondary grounds. It may be confused with a generic concept or with a better-known competitor. It may exist inside the answer, but not occupy a meaningful place within it. And for business, that functional visibility is exactly what matters: not abstract recognition, but participation in the real moment of choice. The change in user behavior makes this problem especially costly. According to McKinsey, roughly half of consumers already use AI-assisted search intentionally, and 44% of those users call it their primary source of information for decision-making [8]. Google reported that AI Overviews had reached more than 2 billion monthly users, while AI Mode had already surpassed 100 million monthly active users in the United States and India [9]. In February 2026, OpenAI reported more than 900 million weekly active ChatGPT users [10]. When interfaces at that scale become the first point of contact for a question, machine invisibility stops being a research curiosity. It becomes a loss of attention share before the click. It is important to emphasize that this is not about AI being “unfair” to brands. Answer systems work differently from classic search. They do not just find documents; they immediately perform an interpretation: which attributes of an entity to treat as primary, which sources to use to validate the answer, which alternatives to name alongside it, how to formulate the category, and what degree of confidence to display. If a brand is not prepared for that environment, it loses not because it lacks a site, but because it lacks a machine-stable form. That stability can be described as a combination of five layers. The first layer is identity: what the company is called, what spelling variants exist, and how the legal name differs from the product name. The second is classification: which category of solutions the brand actually belongs to. The third is properties: which problems it solves, how it differs, and what limitations it has. The fourth is relationships: its products, clients, analogs, partners, geographies, and sectors of application. The fifth is the evidence base: which external sources validate all of the above. When one of these layers is weak, the machine starts completing the picture on the basis of probability rather than clear knowledge. And this is exactly where strong brands turn out, unexpectedly, to be vulnerable: recognition substitutes for precision, and reputation substitutes for structural clarity. ## What this changes in brand management For companies, this leads to a difficult but useful conclusion. In the age of AI, it is not enough simply to be visible; you have to be read correctly. It is not enough to accumulate mentions; you have to build a coherent entity contour. It is not enough to dominate your own channels; you have to be present in the network of validation on which answer systems rely. It is not enough to formulate positioning once; you have to test whether that positioning holds in machine retellings. That is why, today, a strong brand is no longer only a cultural and market object, but also an object of machine knowledge. And the sooner a company takes that seriously, the less tempted it will be to look for a single magic button. The problem of visibility in AI almost never comes down to one button. It almost always comes down to how well the brand has been assembled as an entity — for people, for the web, and for the machines that now increasingly mediate between them. ## What seems well established What is well established is that modern answer systems do not work like a static reference book: they assemble an answer from parametric memory, current context, and external sources. That is why a brand may be known to the system by name and still fail to participate in the answer at the moment of choice. ## What remains uncertain or platform-dependent What is less firmly established is the exact share of brands affected by this kind of invisibility, and whether there is a single set of risk factors across all platforms. The scale of the problem depends on the industry, the language, the type of prompt, and the extent to which the system relies on web retrieval at a given moment. ## Practical implications for brand work The practical implication of this article is that diagnosis should begin not with the question “do they know us?” but with the question “can they consistently assemble us as the correct entity in the relevant scenario?” ## Sources - [1] Geva M., Schuster R., Berant J., Levy O. Transformer Feed-Forward Layers Are Key-Value Memories. EMNLP, 2021. https://aclanthology.org/2021.emnlp-main.446/ - [2] Meng K., Bau D., Andonian A., Belinkov Y. Locating and Editing Factual Associations in GPT. NeurIPS, 2022. https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html - [3] Wang M. et al. Knowledge Mechanisms in Large Language Models: A Survey and Perspective. EMNLP Findings, 2024. https://aclanthology.org/2024.findings-emnlp.416/ - [4] Google Search Central. AI Features and Your Website. 2026. https://developers.google.com/search/docs/appearance/ai-features - [5] OpenAI. Introducing ChatGPT Search. 2024. https://openai.com/index/introducing-chatgpt-search/ - [6] Perplexity Help Center. How does Perplexity work? 2026. https://www.perplexity.ai/help-center/en/articles/10352895-how-does-perplexity-work - [7] Wang Y. et al. Factuality of Large Language Models: A Survey. EMNLP, 2024. https://aclanthology.org/2024.emnlp-main.1088.pdf - [8] McKinsey. Winning in the Age of AI Search. 2025. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search - [9] Google. Alphabet Q2 2025 Earnings Call: CEO's Remarks. 2025. https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2025/ - [10] OpenAI. Scaling AI for Everyone. 2026. https://openai.com/index/scaling-ai-for-everyone/ --- Source: https://ansmeter.com/knowledge-base/internal-brand-representation # Article 2. What AI Really “Knows” About a Company: the Brand’s Internal Representation *Overview and methodology article.* **Research question.** What exactly does an answer system “know” about a company if there is no literal brand card anywhere inside the model. **Evidence type.** Work on transformer-model interpretability, studies of factual retrieval, and surveys of knowledge mechanisms in large language models. **Data freshness.** The theoretical part of the article draws on stable academic findings from 2021–2025; the applied observations are aligned with platform conditions as of March 2026. > Inside the model, the brand exists not as a card, but as a probabilistic landscape of associations. Managing visibility does not mean “adding more mentions”; it means making that landscape more stable and internally consistent. ## Why the card metaphor is misleading When executives hear that a language model “knows” their company, an intuitive image almost inevitably appears in the mind: somewhere inside the system there seems to be a brand card with a name, a short description, a set of properties, and several links to the market. The image is convenient, but wrong. A modern answer system stores information about an entity in a form that looks far less like a reference entry and far more like a distributed network of probabilistic connections. A brand does not occupy one neat slot. There are traces in the model parameters, activatable patterns, hidden states in the current computation, and, in search modes, external documents that are blended in at the moment of response. This distinction matters not only for researchers. As long as a company imagines a “brand card,” it tends to look for simple fixes: add more mentions, rewrite the headline, publish one more page of self-description. But if the brand inside the model is structured as a complex system of associations, the task changes. Then the company has to think not only about the quantity of signals, but about how those signals are organized: how stably the name is linked to the category, how clearly the products are differentiated, how consistently the properties are validated, and how easily the model can distinguish your entity from neighboring ones. ## What interpretability research shows Interpretability research over the past several years has gradually made this internal picture less mysterious. Work by Mor Geva and coauthors showed that the feed-forward blocks of transformer architectures often behave like a kind of “key-value” memory: some input text patterns activate others and push the model toward specific lexical continuations [1]. Work by Kevin Meng and colleagues on locating and editing factual associations showed that some facts in autoregressive models can indeed be linked to relatively localizable computational nodes, especially in the middle layers [2]. A later paper by Masaki Sakata and coauthors found that mentions of the same entity tend to form distinguishable clusters in the internal representation space, while information associated with that entity is often concentrated in a compact linear subspace in the model’s early layers [3]. Finally, survey work on knowledge mechanisms in large language models underscores a general conclusion: knowledge in such systems is real, but distributed, fragile, and dependent on the mode of retrieval [4][5]. The simplest way to picture it is this. Inside the model, the brand exists as a probabilistic landscape. On that landscape there are regions where the company name lies close to words such as “analytics,” “security,” “platform,” “forecasting,” “enterprise market,” or, say, “customer experience management.” There are links to known products. There are traces of old press releases. There is proximity to competitors. There are traces of user questions that, in the training data, were often followed by particular kinds of answers. When the model receives a new prompt, it does not “pull out a card.” It traverses that landscape and assembles the most probable interpretation. That is why the question “what does AI know about the company?” is better replaced with another one: “what configuration of connections can AI reconstruct about the company, stably, across different contexts?” That formulation is both more precise and more useful. What matters for business is not the model’s abstract awareness, but its stability. If you ask the system the same thing in ten closely related ways, will it keep assigning the brand to the same category? Will it keep linking it to the same core properties? Will it correctly distinguish the product from the company, the company from the parent structure, and the legal name from the consumer-facing one? Or will each new prompt summon a slightly different entity? ## Probabilistic landscape, vectors, and stable links That stability is clearly visible in vector representations (embeddings), the numerical forms into which words, phrases, and fragments of context are translated. The proximity of two such representations is often measured using cosine similarity: cos(theta) = (x · y) / (||x|| ||y||) Here x and y are two vectors. One may correspond to a set of brand mentions, the other to a feature such as “enterprise analytics” or “low-cost consumer service.” If the cosine is close to one, the vector directions are similar, and the system tends to treat those objects as tightly connected. If the value is low, or if it changes from one context to another, the connection is weak or unstable. A company does not have direct access to such vectors inside closed commercial models. But the logic is still useful: a brand benefits when the important links in its machine representation are not accidental, but repeatable. This also clarifies the nature of typical distortions. If the brand name is ambiguous, the model may pull it too tightly toward a general category and erase its distinctiveness. If the company has several product lines described in different languages, they may fail to cohere into a single family inside the model. If the external environment knows the old version of the brand better than the new one, the model will “remember the past” more stubbornly than marketing would like. If competitors have a sharper and better-validated semantic contour, a prompt about a class of solutions will lead not to your company, but to them. And the reverse is also true: if the brand is systematically present in the language of the market, in independent sources, and in its own clear descriptions, the model is more likely to assemble your company specifically, even if it is not the largest player. ## Three layers of internal representation and a new diagnostic lens The brand’s internal representation can usefully be divided into three layers. The first layer is parametric memory. This is what the model absorbed during training and subsequent tuning: general facts, typical associations, and habitual links between the name and its properties. The second layer is contextual assembly. This is how the brand is reconstructed at answer time from the hidden states of the current dialogue: which words in the user’s prompt activated which parts of the system’s knowledge. The third layer is external reinforcement. In answer and search modes, fresh web pages, documents, and knowledge bases are added here, and they influence the final output [4][6][7]. In practice, it is the interaction among these three layers that determines what the brand will look like in the answer. This architecture explains why many companies misdiagnose the problem. When a brand is not named in the answer, the usual assumption is that “the model does not know us.” Sometimes that is true, but not always. The model may know the company by name and still fail to consider it the best answer to the question. It may remember the product, but fail to connect it to the right use case. It may cite the site correctly, yet rank the importance of properties incorrectly. It may rely on current web sources and, in doing so, override older internal knowledge. In other words, the problem may lie not in the presence of knowledge, but in its configuration. This is especially important for brands that are used to relying on the force of their own communication. Inside answer systems, the winner is not only the one that speaks loudly about itself, but the one about whom a coherent representation can be built. And a coherent representation requires discipline. The name has to be stable. The category has to be clear. The product structure has to be distinguishable. The properties have to be stated directly, not merely implied. External validation has to be diverse and reliable. Only then does the model have a chance not merely to recognize the brand, but to retain it as a stable entity. This leads to another important conclusion. Working on the brand’s internal representation does not reduce to “text optimization.” At bottom, it is work on the company’s epistemic form — that is, the form in which the company exists as knowledge. When the brand is poorly assembled as knowledge, the answer system is forced to fill in the gaps probabilistically. When the brand is well assembled, the probability of distortion falls. In that sense, the modern struggle for visibility is not only a struggle for traffic, but also a struggle for the quality of machine understanding. This perspective is useful for another reason as well: it moves the conversation onto more mature ground. The question is not whether “AI remembers us.” The question is which properties of our brand are extracted stably, which links are lost, which attributes are overweighted, and which do not make it into the answer at all. Those are the questions from which strategy, diagnostics, and substantive work begin. They are what distinguish serious management of machine visibility from a superficial race for random mentions. ## What seems well established We can say with confidence that knowledge in modern language models is distributed and retrieved contextually. It follows that a brand’s stability in answers cannot be reduced to the mere presence of its name in the training material. ## What remains uncertain or platform-dependent What is less firmly established is the exact geometry of that knowledge in closed commercial systems. Academic work reveals the general mechanisms, but we do not have direct access to the internal vectors and assembly rules of each platform. ## Practical implications for brand work For a company, this means shifting from the language of “text optimization” to the language of epistemic form: it needs to monitor which brand properties are extracted stably and which ones fragment or become distorted. ## Sources - [1] Geva M., Schuster R., Berant J., Levy O. Transformer Feed-Forward Layers Are Key-Value Memories. EMNLP, 2021. https://aclanthology.org/2021.emnlp-main.446/ - [2] Meng K., Bau D., Andonian A., Belinkov Y. Locating and Editing Factual Associations in GPT. NeurIPS, 2022. https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html - [3] Sakata M., Yokoi S., Heinzerling B., Ito T., Inui K. On Entity Identification in Language Models. Findings of ACL, 2025. https://aclanthology.org/2025.findings-acl.858.pdf - [4] Wang M. et al. Knowledge Mechanisms in Large Language Models: A Survey and Perspective. EMNLP Findings, 2024. https://aclanthology.org/2024.findings-emnlp.416/ - [5] Wang Y. et al. Factuality of Large Language Models: A Survey. EMNLP, 2024. https://aclanthology.org/2024.emnlp-main.1088.pdf - [6] Yadav I., et al. External Knowledge Integration in Large Language Models: Survey, Methods, Challenges, and Future Directions. Semantic Web Journal, 2025. https://www.semantic-web-journal.net/system/files/swj3835.pdf - [7] Google Search Central. AI Features and Your Website. 2026. https://developers.google.com/search/docs/appearance/ai-features --- Source: https://ansmeter.com/knowledge-base/source-ecology-and-citation # Article 3. Which Sources Shape AI’s Opinion of a Brand — and Why the Website Is Not the Main Character *Overview and analytical article.* **Research question.** Which layers, exactly, shape machine opinion about a brand, and why the company’s own website remains important but ceases to be the sole arbiter. **Evidence type.** Documents from Google, OpenAI, Microsoft, and Perplexity, as well as survey research on retrieval-augmented generation and on the integration of external knowledge. **Data freshness.** The factual material on search mechanics and answer systems is current as of March 2026. > The website is the core of the brand’s self-description, but it does not hold a monopoly on truth. Machine opinion is assembled from owned documentation, search context, external editorial sources, user traces, and structured knowledge. ## The website as the primary source, but not the only arbiter When a brand first encounters the problem of visibility in AI, its first instinct is almost always the same: improve the company website. The instinct is sound, but incomplete. The official website does remain the central carrier of primary information about the company: it is where the brand explains who it is, what it does, how the product works, what the prices are, what the constraints are, which usage scenarios it serves, and what evidence supports its competence. But in the answer environment, the website no longer functions as the sole and uncontested source of truth. It is an important case document, but not the only witness. And the decision about how exactly to restate the brand for the user is increasingly made by AI on the basis of several source types at once. This follows from the architecture of modern systems itself. Google Search Central states explicitly that AI Overviews and AI Mode use a fan-out decomposition of the query across subtopics and data sources, then surface a broader and more diverse set of supporting links than classic web search [1]. In its AI Mode help documentation, Google adds that the system breaks the question into subtopics and simultaneously looks for relevant material for each of them [2]. OpenAI describes ChatGPT Search as a mechanism for producing fast, up-to-date answers grounded in web sources and informed by the context of the entire conversation [3][4]. Perplexity expresses the same idea with maximum directness: the system searches the internet in real time, gathers information from trustworthy sources, and condenses it into a short explanation [5]. In the research literature, this family of practices is commonly described as a combination of the model’s parametric knowledge and generation with external knowledge retrieval (retrieval-augmented generation) [6][7]. If we translate that technical picture into the language of the brand, the conclusion is simple but important. AI’s opinion about a company is built from at least five layers. ## Five layers of the source contour The first layer is the brand’s owned channels. These include the website, documentation, FAQ sections, product descriptions, pricing pages, case studies, public research, the press center, expert blogs, and, in some cases, video transcripts and technical knowledge bases. This layer defines the base thesaurus: what the brand calls itself, which category it places itself in, and which properties it puts in the foreground. If there is already confusion across the brand’s own channels, no amount of external reputation will save it. The machine needs a starting scaffold. The second layer is search and link context. Even when the answer shown to the user looks like a conversation, the logic of search infrastructure is often still operating underneath it. Google reminds us that, to participate in AI features, pages must be indexed and broadly suitable for ordinary search [1]. Put simply, the AI intermediary rarely starts from zero: it relies on the preexisting layer of discovery, indexing, and selection of web documents. That is why the site’s technical accessibility, the quality of its text, and basic search discipline still matter. But they no longer guarantee dominance. They merely get the brand into the game. The third layer is external editorial and industry sources. These include reviews, comparisons, rankings, interviews, analytical materials, trade-media publications, directories, and business profiles. This is usually where the brand gets what it cannot give itself: external validation. If the official website claims that the company is strong in complex enterprise analytics, while independent sources describe it as a niche tool for small business, the answer system has to reconcile those versions. And very often it chooses the version that is better validated and more clearly embedded in the network of links. In answer systems, self-presentation without external verification carries less weight than brands would like. The fourth layer is the user trace. This includes reviews, forum discussions, questions and answers in communities, mentions on social platforms, opinion catalogs, support pages, and, more broadly, the whole living and not always tidy fabric of the internet in which people explain to one another what a product is and how it works. This layer is noisy and unreliable, but it cannot be ignored. It often shapes the language of actual demand. A company may describe itself as a “modular environment for intelligent data management,” while users discuss it as “a convenient service for complex reporting without a heavy implementation.” For AI, that language matters a great deal, because it is the language in which everyday questions are actually phrased. The fifth layer is structured knowledge. This includes entity databases, open knowledge graphs, catalogs, business directories, organization profiles, standardized descriptions, and, sometimes, schema markup on the site itself. Survey work on integrating external knowledge into language models shows that linking AI to knowledge bases and graphs improves the factual accuracy, traceability, and explainability of the answer [6][8]. For the brand, this means that the role of “boring” and formal-looking sources increases. They rarely create a vivid reputation, but they often provide stable identification of the entity. ## Why answer platforms read this environment differently It is precisely the combination of these layers that explains why the website does not become the main character. It may be the main primary source, but not the main arbiter. Answer systems assess not only what the brand claims about itself, but also how that claim is validated, repeated, challenged, or reformulated by other participants in the network. Put more sharply: the website explains what the brand would like to be seen as; the external environment shows what it is actually seen as; and AI tries to assemble a workable compromise between those versions. Several practically important consequences follow from this. First, it is impossible to work seriously on visibility in AI if you limit yourself to the homepage. Even a brilliantly written website does not guarantee that the brand will be named in the answer if external sources either fail to validate its key properties or validate them differently. Second, the official website remains critically important — precisely because it defines the canonical structure of the entity. But its function changes. It must be not only an attractive storefront, but also a reliable point of alignment: a place where AI and humans can see the same name, category, properties, and evidence with equal clarity. Third, the brand has to manage not only its own text, but also the ecosystem of validation: who writes about it and how, which comparisons it appears in, which catalogs and knowledge bases it is present in, where its methodology is represented, and who can independently validate its role in the market. ## From editing the website to managing the entire knowledge contour What matters especially is that different AI platforms read this environment differently. Google relies on its own search infrastructure and AI modes, where indexability and page eligibility for display matter [1][2]. ChatGPT Search brings in web sources either on request or automatically, while taking the dialogue context into account [3][4]. Perplexity emphasizes almost continuous real-time web retrieval and explicit links [5]. Microsoft Copilot likewise describes its answers as grounded in web search and external links [9][10]. For a brand, this means there is no single “source of truth” from which every machine will read the company in the same way. There is a network of sources that each system assembles according to its own logic. That is why a mature strategy begins with a more mature question. Not “how should we describe ourselves better on the website?” but “what set of sources forms machine opinion about us — and where in that set are we strong, and where are we being undermined by noise, absence, or someone else’s interpretation?” Only after that question does content work stop being cosmetic and become knowledge management. This is exactly where the brand’s new role on the internet comes into view. It used to be able to think of the website as the main stage, and everything else as noisy background. Now the picture flips. The site remains the stage, but the performance has long since stopped unfolding only there. The whole internet stages it. And the answer system acts not as a spectator, but as an editor, assembling the final version for the user out of a multitude of voices. In that logic, the winner is not the one that speaks most loudly about itself, but the one whose entity is validated most clearly and consistently across the network. ## What seems well established It is well established that answer systems use not one document and not one type of signal. For a stable presence, a brand needs a set of aligned sources, not merely a strong homepage. ## What remains uncertain or platform-dependent The exact relative importance of each layer — the website, external media, reviews, catalogs, knowledge graphs — varies from platform to platform and is rarely disclosed in full. ## Practical implications for brand work The practical conclusion is straightforward: what must be managed is the entire source contour. An audit of visibility in AI begins with a source map, not with editing a single paragraph on the website. ## Sources - [1] Google Search Central. AI Features and Your Website. 2026. https://developers.google.com/search/docs/appearance/ai-features - [2] Google Search Help. Get AI-Powered Responses with AI Mode in Google Search. 2026. https://support.google.com/websearch/answer/16011537 - [3] OpenAI. Introducing ChatGPT Search. 2024. https://openai.com/index/introducing-chatgpt-search/ - [4] OpenAI Help Center. ChatGPT Search. 2026. https://help.openai.com/en/articles/9237897-chatgpt-search - [5] Perplexity Help Center. How does Perplexity work? 2026. https://www.perplexity.ai/help-center/en/articles/10352895-how-does-perplexity-work - [6] Yu H. et al. Evaluation of Retrieval-Augmented Generation: A Survey. 2024. https://arxiv.org/abs/2405.07437 - [7] Zhao P. et al. Retrieval-Augmented Generation for AI-Generated Content: A Survey. Data Science and Engineering, 2026. https://link.springer.com/article/10.1007/s41019-025-00335-5 - [8] Ibrahim N. et al. A Survey on Augmenting Knowledge Graphs with Large Language Models. Discover Artificial Intelligence, 2024. https://link.springer.com/article/10.1007/s44163-024-00175-8 - [9] Microsoft. Copilot Search in Bing. 2026. https://www.microsoft.com/en-us/bing/copilot-search - [10] Microsoft Support. Understanding Web Search in Microsoft 365 Copilot Chat. 2026. https://support.microsoft.com/en-us/topic/understanding-web-search-in-microsoft-365-copilot-chat-94c45d32-1a77-4f82-8e05-58dfb9afac48 --- Source: https://ansmeter.com/knowledge-base/path-shift-to-ai-mediator # Article 4. From Search Engine to AI Mediator: How the Customer Journey Is Changing *Overview and analytical article.* **Research question.** How does the customer journey change when the search interface turns into an AI mediator that synthesizes an opinion at the very outset and shortens the path to choice. **Evidence type.** Data from McKinsey, Similarweb, and Adobe, along with official statements from Google and OpenAI on the scale of AI mode adoption. **Data freshness.** The current figures and scenario estimates reflect market conditions as of March 2026. > Search is not disappearing, but its visible surface is changing. Users increasingly receive not a list of documents, but a first synthesized answer — and that is where the new struggle for attention, trust, and demand begins. ## Search remains infrastructure, but its interface is changing The current inflection point on the internet is easiest to describe through one simple substitution. Not long ago, users asked the network for a path to information. Now, increasingly, they ask it for an already assembled judgment. In the old logic, a person entered a query, received a list of links, opened several pages, manually compared wording, prices, and signs of reliability, and then built the picture for themselves. In the new logic, a substantial part of that work is delegated to the AI mediator. The user begins not with navigation across documents, but with a conversation: “explain,” “compare,” “what should I choose,” “who is stronger in this category,” “what is the difference,” “which option fits my case.” Only afterward, if the answer proves insufficient, do they move on to sources. This shift does not yet mean the death of search, which is precisely why it is often underestimated. Search as infrastructure remains enormous. But the interaction surface is changing before our eyes. McKinsey writes that about 50% of Google search queries are already accompanied by AI summaries, and that this figure may exceed 75% by 2028 [1]. In 2025, Google reported that AI Overviews had more than 2 billion monthly users, while AI Mode had already surpassed 100 million monthly active users in the United States and India [2]. In February 2026, OpenAI reported more than 900 million weekly active ChatGPT users [3]. In other words, we are no longer dealing with a niche technological experiment, but with a mass shift in the interface through which people access information. It is especially telling that users themselves already perceive this environment as a source of solutions rather than as a curious add-on. According to McKinsey, about half of consumers deliberately use AI-assisted search, and 44% of those users name it as their main and preferred source of information for decision-making — above traditional search, brand websites, and review platforms [1]. This figure matters not only in itself. It means that the point at which the first impression is formed is increasingly located not on the company website, but in an answer the company does not directly control. ## What the new customer journey looks like The classic customer journey in the digital environment could be described as the chain “query -> list of links -> comparison of documents -> shortlist -> action.” In the new environment, an intermediate link increasingly appears between the query and the list of links — a synthesizing interlocutor. The chain now looks different: “question -> AI summary -> follow-up -> preliminary filtering -> possible site visit.” The difference seems subtle, but in practice it changes the entire attention economy. Previously, a brand could lose at the click-through stage, yet still remain present on the results page. Now it can lose earlier — at the level of whether it was included in the machine judgment at all. The speed at which this layer is growing is also confirmed by outbound referral data. Similarweb estimates that in June 2025, AI platforms drove more than 1.13 billion visits to external websites; that was 357% more than a year earlier [4][5]. By its estimate, the average monthly number of visits to generative AI platforms grew 76% year over year, while app downloads rose 319% [4]. Figure 1 shows this one-year surge in referrals to websites. It is important, however, not to fall into statistical hypnosis: according to the same data, Google Search generated about 191 billion referrals in June 2025, still orders of magnitude more [5]. But this is precisely where the central meaning of the shift begins. AI does not have to overtake search in gross traffic volume already today in order to radically change user behavior. It only has to capture the stage of initial judgment. The quality of this new traffic also says a great deal. According to Adobe Analytics, in retail, visitors arriving from generative AI sources viewed 12% more pages per visit and showed a 23% lower bounce rate than users from non-AI sources [6]. In travel, Adobe recorded a 45% lower bounce rate for such traffic [6]. In its documentation for site owners, Google notes a similar effect: clicks from results pages with AI Overviews appear to be higher quality, and users more often spend more time on the site [7]. This is an important detail in the new logic of the customer journey. There may be fewer clicks, but each one increasingly arrives later in the funnel and with higher intent. ## What the data says about the growth and quality of AI-mediated sessions Hence the first serious change for marketing. The traditional top of the funnel no longer lives only in the search results. It partially dissolves into the conversation. The user who previously read five documents and only then formed a brief opinion now receives that opinion at the outset. The AI answer may already name the main market players, note the strengths and weaknesses of solutions, filter out obviously unsuitable options, and highlight pricing and technical constraints. The brand website becomes not the first point of contact, but the place where a choice already underway is refined. The second change is subtler still: for ordinary users, the distinction between “search” and “talking to AI” is rapidly losing meaning. What matters to them is not the technological mode, but the user experience. They see a short answer at the top, can ask a follow-up question, get links, and continue the conversation. In such an environment, the search engine itself becomes an answer system, while chat becomes a form of search. That is why the debate of “search versus AI” now describes reality poorly. What we are seeing, rather, is the dissolution of the old boundary between search, reference, comparison, and consultation. The third change concerns the tempo of decision-making. The AI mediator can shorten the path not only through speed, but through cognitive offloading. It takes on the preliminary comparison, the translation of complex language into simpler terms, the condensation of large amounts of material into a few paragraphs, and the delivery of first judgments. For the person, that is convenient. For the brand, it means that a substantial share of the struggle for attention now takes place earlier and in a more compressed form. If a company fails to make it into that preliminary filtering stage, its chances of recovering later are lower. ## A moderate forecast and the website’s new role To describe the next stage of this transformation, we suggest looking not at the “market share of individual platforms,” but at a broader measure: the share of informational sessions in which the first answer ultimately delivered to the user is meaningfully mediated by an AI layer. This includes AI Overviews, AI Mode, standalone chat interfaces, and other answer surfaces where the user receives a first interpretation before moving to documents. Figure 2 presents an editorial forecast of this share over the next 1, 3, and 5 years. In the base case, we estimate it at roughly 18% for 2026, 25% for 2027, 39% for 2029, and 54% for 2031. This is not an official market metric, but an analytical estimate built from a combination of public signals: the scale of AI Overviews, the size of the ChatGPT audience, the growth of AI referrals, and the behavioral-shift data published by McKinsey, Google, OpenAI, Similarweb, and Adobe [1][2][3][4][6]. This forecast is intentionally moderate. It does not assume that classical search will disappear. On the contrary, search will remain the deep infrastructure of retrieval and verification. But the user-facing surface will increasingly look not like a list of links, but like an answer with the option to go deeper. One year out, this most likely means an acceleration of a process already underway: AI will even more often become the first point of contact with a question, but a substantial share of site traffic will still come from classical search, direct visits, and other channels. Three years out, the distinction between search and AI will become much less meaningful for the mass user: they will remember not the mode, but the convenience. Five years out, the most probable world is one in which, in a substantial share of informational sessions, the market’s first filter is precisely the answer layer. All of this also changes the role of the brand website. Previously, the site aspired to be the stage of first contact. Now that role increasingly passes to the AI mediator. The site becomes a place for checking, deepening, confirmation, comparison of details, commercial contact, and action. That does not make it less important. But it does make it part of a longer and more complex chain. In that chain, the brand has to be understandable not only to the search index and not only to the person who has already arrived on the site, but also to the answer system that decides whether it is worth sending the user there at all. That is why the current paradigm shift looks not like an instantaneous revolution, but like the gradual disappearance of the gaps between channels. Search, reference, comparison, recommendation, and consultation are being pulled into a single point. That point is the AI mediator. And the struggle for visibility, trust, and demand begins even before the user reaches the brand page — in the structure of the first answer. ## What seems well established It is clearly established that the AI layer is already embedded in mass search and comparison scenarios, and that part of traffic and decision-making is shifting toward answer interfaces even before the click to the brand website. ## What remains uncertain or platform-dependent Any precise forecast of the share of AI-mediated sessions over a multi-year horizon remains scenario-based. The pace depends on the country, the query type, user trust, and the speed with which new modes are integrated into familiar platforms. ## Practical implications for brand work The implication for brands is that the site ceases to be the only stage of first contact. It becomes a place for verification, deeper exploration, and action, while the first filter of selection increasingly happens outside it. ## Sources - [1] McKinsey. Winning in the Age of AI Search. 2025. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search - [2] Google. Alphabet Q2 2025 Earnings Call: CEO's Remarks. 2025. https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2025/ - [3] OpenAI. Scaling AI for Everyone. 2026. https://openai.com/index/scaling-ai-for-everyone/ - [4] Similarweb. AI Discovery Surges: Similarweb's 2025 Generative AI Report Says. 2025. https://ir.similarweb.com/news-events/press-releases/detail/138/ai-discovery-surges-similarwebs-2025-generative-ai-report-says - [5] Similarweb. AI Referral Traffic Winners by Industry. 2025. https://www.similarweb.com/blog/insights/ai-news/ai-referral-traffic-winners/ - [6] Adobe. Adobe Analytics: Traffic to U.S. Retail Websites from Generative AI Sources Jumps 1,200 Percent. 2025. https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent - [7] Google Search Central Blog. Top Ways to Ensure Your Content Performs Well in Google's AI Experiences on Search. 2025. https://developers.google.com/search/blog/2025/05/succeeding-in-ai-search --- Source: https://ansmeter.com/knowledge-base/market-landscape-ai-visibility # Article 5. What the market offers to increase visibility in AI—and where the hidden costs of these approaches lie *Market overview with analytical commentary.* **Research question.** What the market for AI visibility tools can already do, where its hidden limitations lie, and why polished dashboards do not replace a targeted strategy for a specific brand. **Evidence type.** Official product pages from Adobe, Similarweb, Brandlight, Scrunch, Semrush, and other players; publicly available information on pricing and sales models. **Data freshness.** Product descriptions and pricing signals were checked against open materials available in March 2026. > The market has already learned to observe AI visibility quite well, but it is much weaker at turning observation into a targeted action program. Its typical constraints are high cost, enterprise orientation, and a machine-centered view of the problem. ## From manual checks to observation platforms The market for AI visibility solutions has matured surprisingly quickly. Not long ago, teams were checking ChatGPT or Perplexity responses manually: they asked a few questions, took screenshots, and argued over whether the result should count as a signal or a coincidence. Today there is already an entire product category promising to monitor brand mentions, compare visibility against competitors, track citations, show shifts in sentiment, surface content opportunities, and even suggest technical fixes. This is an important and useful stage in the market’s maturation. But it is also the stage at which it becomes especially easy to confuse category maturity with the problem itself having been solved. If you look closely at what the market leaders are actually selling, a fairly coherent logic emerges. Most platforms offer brands three core modules. The first is observation: how often different AI platforms mention you, in what wording, against which competitors, and with what sentiment. The second is causal analysis: which pages are being cited, which prompts do or do not produce mentions, where visibility collapses, and which technical or content gaps are getting in the way. The third is action: recommendations for revisions, content suggestions, diagnostics of technical issues, and sometimes dedicated solutions for delivering a more machine-readable version of the site. At the level of the promise, it all looks convincing. At the level of real implementation, it becomes more complicated. Adobe frames the issue in the language of enterprise marketing. Adobe LLM Optimizer promises brands the ability to manage how they appear in AI search, measure AI traffic, track “share of voice,” and receive prescriptive recommendations, including automated fixes [1][2]. Similarweb structures its offer around AI Search Intelligence: brand visibility, prompt analysis, citation analysis, sentiment, and actual traffic from AI platforms. In its documentation, the company explicitly emphasizes that the module shows how often a brand appears in language model responses and which sites influence that presence [3]. Profound states the task even more directly: to make sure a brand is named and recommended in conversations with AI. Its site emphasizes monitoring of answer systems, agent analytics, and growth in visibility within responses [4]. Brandlight openly describes itself as an “AI visibility platform for enterprise brands,” emphasizes work with Fortune 500 companies, and promises a unified view of how a brand is represented in AI search [5]. seoClarity promotes ArcAI as an “enterprise” framework for analyzing and fixing AI search visibility, where data is translated into prioritized recommendations for the team [6]. Scrunch, for its part, combines monitoring, causal analysis, and a specialized content delivery layer for AI agents through its own Agent Experience Platform—that is, a special “lightweight” version of the site for machine reading [7]. This is impressive in itself. In a short time, the industry has gone from scattered observations to systematic instrumentation. Brands now have dashboards where they can see prompts, citations, visibility dynamics, the relationship between topics and responses, and sometimes even separate signals showing how AI agents are crawling the site. For large teams, this is a major relief: the topic has moved out of the realm of intuition and become measurable. The market deserves credit for exactly that. It legitimized the problem itself. ## What exactly the market leaders offer But this is also where the hidden costs begin. The first is the clear enterprise orientation of the leaders. You can see it not so much in the marketing language as in the sales model itself. On its pricing page, Profound talks about “custom enterprise pricing” and explicitly describes the platform as a solution for global brands [8]. Brandlight sells itself as an enterprise platform for the largest companies and routes buyers toward a demo rather than a transparent product tier [5]. Adobe positions LLM Optimizer as a solution for business and large digital marketing teams [1][2]. seoClarity consistently uses the language of enterprise-grade infrastructure and cross-team coordination [6]. Even where entry-level product tiers do exist, serious use cases almost always push the buyer toward more expensive plans, additional licenses, and internal approval processes. The second hidden cost is not just price as such, but total cost of ownership. Similarweb, for example, offers a self-serve AI Search Intelligence plan for $99 and an expanded tier for $399, but the task set itself quickly pushes the user toward a broader data stack and eventually into a sales conversation [3]. Semrush positions its AI Visibility Toolkit much closer to the mid-market, but its documentation separately notes extra charges for additional user licenses and for new domains or locations [9]. Even comparatively “lightweight” solutions almost inevitably become more expensive once a company wants to move beyond experimentation and start working systematically. And if a brand operates across several markets, with multiple sites, products, and teams, cost stops being a question of a single subscription and becomes a question of organizational architecture. ## Hidden costs: expense, enterprise bias, and an incomplete picture The third hidden cost is the machine-centered angle of view. Almost all of the stronger platforms are very good at answering the question, “What is happening in answer systems?” They show mentions, presence share, citations, sources, sentiment, and sometimes technical signals of site crawling. But they are much weaker at answering another question: “What language is the market itself using to express the problem, and why does the brand fail to appear in that language?” Those are not the same question. You can measure prompts flawlessly and still fail to understand that the brand describes itself in language users do not actually use. You can see citations and still fail to understand that the model does not consider the company relevant not for technical reasons, but because the categorical frame itself has been set incorrectly. Many solutions still provide a powerful instrument here, but not always an interpretation. The fourth hidden cost is the illusion of completeness. The more polished the dashboard, the easier it is to forget that it shows only the part of reality that could be formalized. In AI visibility, that is especially dangerous. Answer systems are stochastic, platforms change quickly, sources are blended in different ways, and human language rarely fits into a neat set of trackable prompts. When a dashboard shows that a brand appears in 18% of responses, the temptation is strong to treat that number as an almost physical fact. But without qualitative interpretation, that number can be a trap. It does not tell you in which scenarios the brand is critically invisible, which model errors are more costly than others, where the problem lies in the site, where it lies in the external trust contour, and where it lies in how the question itself has been framed. The fifth hidden cost is dependence on the client’s internal resources. By design, the best platforms assume that the company already has people in place to execute the recommendations. You need people who will rewrite pages, fix the technical structure, build relationships with external sources, rework terminology, rethink comparison pages, change the content architecture, and measure the effect. For the largest brands, that is natural. For many mid-sized companies and for niche B2B players, it is far less obvious. As a result, the tool is purchased to make the problem visible, but the resources to solve it do not always follow. ## Why the next step is interpretation and targeted recommendations To avoid oversimplifying the picture, it is important to say what is working well. Each of the market leaders has a real strength. Adobe and Similarweb know how to speak to management in the language of traffic impact and business metrics [1][3]. Profound and Brandlight package the issue effectively as a brand management problem in an AI environment [4][5]. seoClarity focuses on translating data into executable recommendations for large teams [6]. Scrunch is trying to solve a problem that remains rare in the market: not only measuring, but reworking the site’s mode of presentation for machines [7]. Semrush is easier to understand than many enterprise players for the mid-market [9]. So the problem is not that there are no solutions. The problem is that almost all of the best solutions are either expensive, require a mature internal team, or remain too concentrated on the machine layer and not sufficiently sensitive to the human language of choice. That is exactly why the market is in an intermediate stage. It has already learned to observe AI visibility reasonably well, but it has not yet fully learned how to turn that observation into a targeted strategy for a specific brand. And if we look at the situation soberly, the next wave of value will not be created where the dashboard becomes even brighter, but where diagnostics more precisely connect machine signals to the real structure of demand: the user’s language, the structure of the category, the set of external confirmations, and the client’s own constraints. For large corporations, today’s market leaders are already genuinely useful. For the rest of the market, their promise is often harder to execute than it appears in the initial demo. And that is perhaps the clearest sign of the moment. The industry has built good instruments. But the real work—interpretation, prioritization, and point-by-point change to the brand’s machine image—remains far less automated than the sellers of those instruments would like. ## What seems well established It is safe to say that the market’s tools already capture mentions, citations, prompts, sources, and presence share across multiple platforms. It is equally clear that many solutions are sold in the logic of the large enterprise client. ## What remains uncertain or platform-dependent What is less certain is how quickly the current leaders will be able to move from a general observation dashboard to genuinely personalized recommendations by brand, category, and the language of prompts. ## Practical implications for brand work The practical conclusion is that even a good tool is not, by itself, a strategy. For most companies, value emerges only when the data is converted into priorities, sequencing, and context-specific recommendations. ## Sources - [1] Adobe. Adobe LLM Optimizer. 2025. https://business.adobe.com/products/llm-optimizer.html - [2] Adobe Experience League. LLM Optimizer Overview. 2025. https://experienceleague.adobe.com/en/docs/llm-optimizer/using/essentials/overview - [3] Similarweb. AI Search Intelligence. 2026. https://www.similarweb.com/corp/search/gen-ai-intelligence/ - [4] Profound. Optimize Your Brand's Visibility in AI Search. 2026. https://www.tryprofound.com/ - [5] Brandlight. AI Visibility Platform for Enterprise Brands. 2026. https://www.brandlight.ai/ - [6] seoClarity. ArcAI Insights. 2026. https://www.seoclarity.net/ai-seo/ai-search-insights - [7] Scrunch. The AI Customer Experience Platform. 2026. https://scrunch.com/ - [8] Profound. Pricing. 2026. https://www.tryprofound.com/pricing - [9] Semrush. AI Visibility Toolkit. 2026. https://www.semrush.com/kb/1493-ai-visibility-toolkit --- Source: https://ansmeter.com/knowledge-base/economics-of-invisibility # Article 6. The Economics of Invisibility: How a Company Loses Demand Before the First Click *Overview and economic analysis article.* **Research question.** How can the problem of AI invisibility be translated from an abstract conversation about traffic into the language of early economic losses and manageable metrics. **Evidence type.** Market data on changing user behavior, studies on declining outbound clicks, platform documents, and the author’s illustrative calculation model. **Data freshness.** The market evidence refers to 2025–2026; the calculation model is illustrative and intended for internal estimation. > The main loss in the new environment occurs not when a brand fails to get a click, but when it fails to make the shortlist and loses the right to participate in shaping choice before the user ever reaches the site. ## Loss before the click as a new category of loss For a long time, digital marketing operated within a relatively comfortable logic. First came an impression. Then a click. Then on-site behavior. Then a lead or purchase. Every intermediate loss was meant to fit into the funnel and be counted. That model was never perfect, but it was clear. And that is precisely why the shift to AI mediators has been so disorienting for many companies. In the new environment, a noticeable share of selection happens before the user ever reaches a site. The user asks a question, receives an initial judgment, refines it, and only after that — if they still see the need — opens one or two sources. At that point, a brand can lose demand without ever losing a formal impression in the classical sense. It simply will not be included in the machine’s preliminary selection. This is the new economics of invisibility. Its main effect is that demand leaks away not only through missing traffic, but through missing participation in the formation of choice itself. When an answer system answers the question “which platforms are suitable for analytics of user signals?”, it performs several actions that were previously carried out by the user. It defines the boundaries of the category. It chooses whom to place side by side. It suggests the language of comparison. It sometimes sets the price frame in advance. It filters out options it considers weak or irrelevant. And all of this happens before the brand gets a chance to explain itself on its own territory. That is why the problem is no longer reducible to a decline in clicks. The losses become earlier and deeper. A brand can lose its place on the shortlist. It can be mentioned, but in the wrong category. It can appear in the answer without the decisive attribute. It can lose to a competitor not because that competitor has a better website, but because the machine judged its description to be clearer and better substantiated. And the reverse is also true: when a brand is woven into the answer correctly, the site often receives a higher-quality visitor — someone who has already passed the coarse filtering stage and has come to clarify details. Public data confirms that this shift is already affecting the real economics of channels. McKinsey writes that unprepared brands may see traffic from traditional search channels decline by 20–50% as decision-making shifts to AI platforms before the click [1]. Google, meanwhile, notes that clicks from pages with AI Overviews tend to be “higher quality”: users are more likely to spend longer on the site [2]. Adobe Analytics records a similar pattern: in retail, visitors from generative AI sources view 12% more pages and show a 23% lower bounce rate; in travel, the bounce rate is 45% lower [3]. In other words, the new environment makes traffic simultaneously smaller in volume and richer in meaning. To lose it is to lose not a casual pageview, but often a more mature intention. ## A simple calculation model and an illustrative example To understand the scale of the problem, it is useful to introduce a simple calculation model. Let the economic loss from invisibility in AI be estimated as follows: P ≈ N × d × v × c × m where N is the number of relevant informational sessions in your category over the period, d is the share of sessions in which the first answer is already materially mediated by AI, v is the probability that your brand will not be included in the answer or will be included in a weakened form, c is the probability that proper inclusion of the brand could have led to a commercially meaningful step, and m is the average margin or value of such a step. The formula is rough, but useful. Its strength lies not in mathematical subtlety, but in forcing us to see the loss before the click. In the classical logic, many companies count only the actual missed visit. In the new logic, they also need to count the lost probability of consideration. And that already requires a different kind of managerial thinking. Take an illustrative example. Suppose that in a niche, 200,000 highly relevant informational sessions occur per quarter. Suppose that the share of AI-mediated sessions in this niche has already reached 25%. Suppose that, according to an audit, the brand is absent or weakly represented in 60% of key questions. Suppose that only 3% of correct inclusions ultimately lead to a product demo, a lead, or another valuable action. And suppose that the average gross value of such an action for the company is $1,200. Then the expected quarterly loss would look roughly like this: 200,000 × 0.25 × 0.60 × 0.03 × 1,200 = 1,080,000 Of course, this is not a precise accounting calculation. But it shows the main point: the price of machine invisibility is easily measured not in “missed mentions,” but in six- and seven-figure sums. And often that happens before the marketing team has time to notice any visible collapse in web analytics. ## Five economic mechanisms of invisibility Why does this happen? Because the AI mediator affects several economic mechanisms at once. The first mechanism is the contraction of the set of alternatives under consideration. A person who previously opened five to seven links and kept a broad set of options in mind may now receive a shortlist of three or four names from the system. If your brand is not on it, you have lost not a click, but participation in the contest to make the shortlist. The second mechanism is the shift in value toward later and higher-value visits. When Google and Adobe say that referrals from AI answers are higher quality [2][3], that means the early filtering has already happened. Consequently, every site visit that gets through becomes more valuable. But that also makes invisibility more painful: the brand is losing not “top-of-funnel noise,” but a potentially more primed visitor. The third mechanism is the substitution of the comparison frame. If the answer system describes your market in a language that does not favor your positioning, the brand starts losing before the facts are even discussed. For example, a complex enterprise solution may be described as a “heavy and expensive tool,” rather than as a “precise system for high-value scenarios.” A consumer service may be framed as “affordable but limited.” In both cases, the company loses not only its place, but the frame through which it is perceived. The fourth mechanism is the growth in the cost of compensation. When a brand does not receive organic participation in answer systems, it has to compensate through paid channels, work more aggressively on demand interception, expand the sales team, or lower the price in order to break into the market. In other words, invisibility in AI rarely remains a purely “informational” problem. It almost always turns into a problem of acquisition cost. The fifth mechanism is the cumulative effect of erroneous knowledge. If a system regularly cites outdated or inaccurate attributes about a company, this affects not one user, but many micro-scenarios of choice. The reputational damage accumulates slowly but stubbornly: the brand ends up not where it should be, and not as what it believes itself to be. ## How to broaden demand metrics Against this backdrop, it is especially important not to fall into the opposite extreme — the panicked belief that every click from classical search is now doomed to disappear. Similarweb shows that the gross volume of referrals from Google still vastly exceeds referrals from AI platforms [4]. But that is precisely why the moment is so important: companies still have time to rebuild while the old infrastructure has not disappeared and the new one has already become a point of preliminary filtering. For practice, two conclusions follow. First, digital demand metrics need to be broadened. It is no longer enough to track rankings, organic traffic, and post-click conversion. Brands need to measure separately the brand’s participation in answers, the quality of that participation, its role on the shortlist, the correctness of the machine description, and the sources from which that description is built. Second, the economics of AI visibility require a tailored assessment. A universal figure such as “share of mentions” says almost nothing without an understanding of the category, the cost of error, the length of the sales cycle, and the value of a late-stage visit. In essence, we are entering an environment where part of marketing value arises at the moment when the user has not yet clicked anything. That is an unusual thought for the old web, but a completely natural one for the era of answer systems. And that is precisely why companies that learn to measure invisibility before the click will gain an important strategic advantage. They will stop treating AI answers as curious external noise and begin to see them for what they have already become: a new layer of demand distribution. ## What seems well established It is reasonably clear that the AI mediator can redistribute attention before the click and narrow the set of alternatives under consideration. As a result, the economic weight of a brand’s early participation in the answer changes. ## What remains uncertain or platform-dependent Any monetary estimate of losses requires hypotheses about the share of AI-mediated sessions, the probability of influencing choice, and the average value of including the brand on the shortlist. These parameters need to be calibrated on a company’s own data. ## Practical implications for brand work For the team, this means the need to measure not only visits and post-visit conversions, but also the lost probability of consideration: invisibility before the click is gradually becoming a separate line item in the cost of growth. ## Sources - [1] McKinsey. Winning in the Age of AI Search. 2025. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search - [2] Google Search Central Blog. Top Ways to Ensure Your Content Performs Well in Google's AI Experiences on Search. 2025. https://developers.google.com/search/blog/2025/05/succeeding-in-ai-search - [3] Adobe. Adobe Analytics: Traffic to U.S. Retail Websites from Generative AI Sources Jumps 1,200 Percent. 2025. https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent - [4] Similarweb. AI Referral Traffic Winners by Industry. 2025. https://www.similarweb.com/blog/insights/ai-news/ai-referral-traffic-winners/ - [5] Google. Alphabet Q2 2025 Earnings Call: CEO's Remarks. 2025. https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2025/ --- Source: https://ansmeter.com/knowledge-base/mention-citation-influence # Article 7. Mention, citation, and influence: three levels of brand presence in AI answers *Research article.* **Research question.** Why brand presence in AI answers needs to be broken down into mention, citation, and influence, rather than reduced to the bare fact that the name appears at all. **Evidence type.** Bing Webmaster Tools and Google Search documentation, along with research on citations, trust, and the influence of AI summaries on traffic. **Data freshness.** The methodological propositions are aligned with public tools and research as of March 2026. > Mention signals visibility, citation signals what the answer rests on, and influence signals a brand’s ability to shape the architecture of interpretation itself. Without that distinction, any presence metric remains too coarse. ## Why simple presence is no longer enough In the era of classic search, talking about brand visibility was relatively straightforward. A company could look at the results page, measure its position, count clicks, and draw conclusions about awareness. Answer systems have broken that convenient linearity. A brand can now appear in an answer and still receive no visit. It can be cited and still remain secondary. It may not be named directly at all, yet its data, phrasing, or arguments may in practice determine the final synthesis. For companies, that means the old question — “do they see us?” — no longer works on its own. In its place comes a more complex one: “what role do we occupy in the machine-generated answer — a mere name, a validated source, or a force shaping the decision itself?” ## Mention, citation, and influence as three distinct layers That is precisely why it makes sense to distinguish at least three levels of presence in the answer environment: mention, citation, and influence. Mention is the simplest form of visibility. The brand is named in the answer: the system lists market participants, compares options, describes a category, and includes the company among the possible answers. For communication purposes, that is already useful. A user who may never have heard of the brand before at least encounters its name. But from the standpoint of business impact, it is not enough. A mention can be accidental, superficial, and even unfavorable. It says nothing about the brand’s status in the answer or whether the system trusts its data. The next level is citation. Here the brand or sources associated with it become part of the answer’s evidentiary support. At the beginning of 2026, Microsoft introduced a dedicated AI Performance panel in Bing Webmaster Tools that shows when a site is cited in AI answers, which pages are cited most often, and which “grounding” phrases the system uses to find that content [1]. But Microsoft itself explicitly notes an important boundary: citation frequency is not the same as either page importance or that page’s role in a particular answer [1]. In other words, citation is a step forward from a simple mention, but it still does not exhaust the question of influence. A brand may appear frequently in the list of sources while remaining only one background support among many for a broader, depersonalized summary. The deepest level is influence. This is the brand’s ability to determine the architecture of the answer: the comparison frame, the set of criteria, the language used to describe the category, and the understanding of a solution’s strengths and weaknesses. Influence is possible even when the user never clicks a link and the brand is not always visible on the surface. If, in answering a question about how to choose a solution, the system reasons in the same categories the brand has spent years developing in its research and reproduces the same logic of proof, then the brand is already influencing the user’s decision — even if that act of influence never translates into the familiar metric of website traffic. ## How this changes the economics of attention and trust This distinction matters not only in theory. It changes the economics of attention itself. Google explicitly says that AI search features are included in overall search traffic, yet users are asking longer and more specific questions, and clicks from pages with AI Overviews turn out to be “higher quality,” meaning they are more often associated with deeper on-site engagement [2][3]. Microsoft, for its part, reports that the user journey in Copilot-assisted scenarios is on average one-third shorter, while rates of high-intent actions are noticeably higher than in traditional search [4]. That means something simple: part of the decision is now happening before the click. Measuring only site visits therefore means measuring only the tail end of the process, not its causal center. The research literature confirms that citations and links in answer systems serve not only a navigational function but also a psychological one. In the SourceBench paper, the authors state directly that the quality of web sources determines an answer’s “groundedness,” and that the presence of links meaningfully increases user trust, even when users do not go on to verify those sources themselves [5]. Search Arena arrives at a similar conclusion from another direction: on average, users prefer answers with more citations, and this preference is especially strong in recommendation and information-synthesis scenarios [6]. But there is a subtle paradox here. If a link increases trust by itself, then citation becomes not only a channel of verification but also an element of rhetoric. At that point, it is no longer enough for a brand simply to “be on the internet”; it needs to become precisely the kind of source the system benefits from relying on in order to produce a convincing answer. Against that backdrop, the distinction between mention, citation, and influence becomes practically measurable. Mention can be counted as the share of answers in which the brand is named explicitly. Citation can be counted as the share of answers in which the brand’s website or major external sources about it are linked. Influence is harder: it has to be inferred indirectly, through semantic proximity between the final answer and the knowledge the brand has contributed to the source contour. In applied work, this can at least be described conceptually through a simple formula: P = aU + bC + cV where P is the integrated brand presence in the answer environment, U is the share of mentions, C is the share of citations, V is the share of cases of semantic influence, and the coefficients a, b, and c reflect the business value of each level for a given industry. For a media brand, the coefficient on mention may be higher, because publicity itself already has value. For a complex B2B solution, by contrast, the weight of influence will be higher: the brand wins not when it is merely named, but when its expertise determines the criteria of choice and the frame of trust. The formula does not pretend to be a strict academic indicator; its purpose is different. It forces marketing and analytics teams to stop thinking in one dimension. ## What and how to measure in your own database This is exactly where a new problem emerges for business. Many companies are still inclined to celebrate the mere fact that their name appears in AI answers. But a mention may be almost empty. The brand is named among ten alternatives — and that is all. The user receives no reason to see it as trustworthy, no validating sources, no sense of what distinguishes it from the rest. In the language of classic advertising, that is like a logo flashing on screen without any explanation of the product’s role. Citation is better: it creates the appearance of verifiability. But there is a risk here too. The system may cite not the brand’s strongest page, but a random catalog, an outdated review, or a weak product card. In that case, the citation is present, but control is not. The most valuable layer — influence — takes the longest to achieve, because it relies not on a single document but on a coordinated knowledge ecosystem. To influence an answer, a brand must do more than publish a product description. It must anchor its formulations across several source types: its own materials, independent reviews, comparative articles, industry discussions, machine-readable catalogs, and, where possible, reliable external validation. Only then does the answer system begin to perceive the brand not as one name among others, but as an entity around which a stable and reproducible contour of meaning has already formed. New data on user behavior reinforce that conclusion. Research from the Pew Research Center shows that when an AI summary appears, users click ordinary search results noticeably less often than when no summary is present, while links inside the summary itself are clicked only in a small share of cases [7]. Another recent paper using Wikipedia data provides early causal evidence that AI Overviews can reduce traffic even for sources that are frequently cited inside the summary itself [8]. That means citation and traffic have definitively diverged. A brand can be important to the construction of the answer and still fail to receive the traffic it might expect. But it does not follow that citation is useless. On the contrary, citation becomes the intermediate bridge between the brand’s text and its influence on the decision. For Ansmeter, this opens an important research perspective. If you are building your own observation database, it is worth recording not only the binary feature “the brand was / was not in the answer,” but at least three fields: whether it was named, whether it or its surrounding source contour was cited, and whether it determined the semantic structure of the answer. At first glance, that last field may seem subjective. But over time it can be normalized quite well: you can look at whose comparison criteria were used, whose terms are repeated, where the key definitions came from, and whether the logic of the synthesis matches the logic of the brand’s own materials. From this follows the main practical conclusion. In the answer environment, a brand wins not when it is merely visible, but when it becomes intellectually necessary to the answer. Mention gives presence. Citation gives the right to trust. Influence gives participation in the decision. Those who learn to distinguish among these three levels will see the real structure of the new visibility. Those who continue to measure everything by the click alone will be looking at a new map with old eyes — and will inevitably underestimate where demand is now actually being created. ## What seems well established It is possible to state with confidence that mention, citation, and influence are not the same thing. A brand may appear in an answer without being its evidentiary support, let alone defining the frame of choice. ## What remains uncertain or platform-dependent What is less firmly established are the universal weights for an integrated presence index: different industries and different question types make the contribution of each layer unequal. ## Practical implications for brand work The practical meaning here is that teams should stop celebrating the bare fact that the name appeared and start tracking where the brand becomes support for the answer — and where it remains only an accidental passenger. ## Sources - [1] Microsoft Bing Webmaster Tools. Introducing AI Performance in Bing Webmaster Tools Public Preview. 2026. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview - [2] Google Search Central. AI Features and Your Website. 2025-2026. https://developers.google.com/search/docs/appearance/ai-features - [3] Google Search Central Blog. Top ways to ensure your content performs well in Google's AI experiences on Search. 2025. https://developers.google.com/search/blog/2025/05/succeeding-in-ai-search - [4] Microsoft Bing Webmaster Blog. How AI Search Is Changing the Way Conversions are Measured. 2025. https://blogs.bing.com/webmaster/November-2025/How-AI-Search-Is-Changing%E2%80%AFthe%E2%80%AFWay%E2%80%AFConversions%E2%80%AFare-Measured - [5] Zhang Y. et al. SourceBench: Can AI Answers Reference Quality Web Sources? 2026. https://arxiv.org/abs/2602.16942 - [6] Search Arena: Analyzing Search-Augmented LLMs. 2025. https://arxiv.org/abs/2506.05334 - [7] Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results. 2025. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/ - [8] Yu, L., Yoganarasimhan, H., and Zuo, L. Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia. 2026. https://arxiv.org/abs/2602.18455 --- Source: https://ansmeter.com/knowledge-base/answer-bubbles # Article 8. “Answer bubble”: why the same brand looks different in ChatGPT, Google, Copilot, and other systems *Research article.* **Research question.** Why the same query produces different versions of a brand across different systems, and why that is dangerous for diagnosis and strategy. **Evidence type.** Recent academic papers on “answer bubbles,” comparisons between web search and generative answers, and platform documents on the mechanics of AI modes. **Data freshness.** This article draws on research and official documents from 2025–2026. > There is no single AI visibility. There are several different answer worlds, in which a brand can be assembled from different sources, in different words, and with different degrees of trust. ## The illusion of unified AI visibility One of the most dangerous illusions in the new market is the belief that there is such a thing as a single, unified “visibility in AI.” A brand asks a question in one system, sees the answer, and draws a sweeping conclusion: either they see us or they do not. But in reality, the answer environment has already split into several autonomous ecosystems, each constructing its own version of internet reality. That is why the same brand can appear as an established leader in one interface, a debatable option in another, and an almost invisible entity in a third. Not because the brand changed overnight, but because the very machinery of selecting, synthesizing, and presenting knowledge changed. The concept of an “answer bubble” is appropriate here. In spring 2026, the researchers behind the Answer Bubbles paper analyzed 11,000 real-world queries across several systems and showed that AI answer environments display pronounced biases in source selection, in the language of the summary, and in how citations relate to the final claim [1]. What matters most is that the authors identified not merely differences in answer quality, but structurally different information realities. The same queries lead to different source sets, different tones of confidence, and different levels of visibility for particular document types. The study also shows that once search is added, systems reduce the number of uncertainty markers — in plain terms, they begin to sound more confident — while simultaneously reinforcing their own source-selection biases [1]. These are no longer simple stylistic variations; they are differences in the design of the very window through which the user sees the market. ## What an “answer bubble” is made of Why does this happen? The first reason is that systems rely on different search and retrieval infrastructures. Google explicitly explains that AI Overviews and AI Mode use query fan-out across subtopics and data sources — which is the company’s own term for the process — and may show a broader set of supporting links than classic search [2]. But Google also notes that AI Mode and AI Overviews may use different models and techniques, which means the set of answers and links can differ even within the same ecosystem [2]. That is an important nuance. The difference between systems does not run only along the line of “Google versus everyone else,” but also inside each platform, between its different answer modes. The second reason is the difference in models’ parametric memory — that is, the knowledge absorbed before the specific query was ever asked. The paper Navigating the Shift emphasizes that the divergence between traditional search and generative answers is driven not only by current web retrieval, but also by the model’s pretraining, which continues to shape the logic by which sources are selected and interpreted [3]. For a brand, that implies an unpleasant but sobering fact: its presence on the internet does not yet guarantee that all systems will read that presence in the same way. One system leans more heavily on live search and fresh documents, another on pre-learned category patterns, and a third on some blend of the two. The third reason is different source preferences. Answer Bubbles shows that generative summaries disproportionately include Wikipedia and longer texts, while social sources and negatively framed materials are, by contrast, underrepresented [1]. The Rise of AI Search adds another layer to that picture: on average, AI search surfaces less of the web’s “long tail,” links more often to the largest sites, and in general offers less answer diversity than classic search [4]. For the market, this means that different systems do not simply find different documents. They answer a different prior question: what kind of source is worthy of becoming part of the public version of reality at all? The fourth reason is different interface and policy choices. In the already mentioned paper The Rise of AI Search, the authors show that the appearance of an AI answer itself depends on query type: question-like queries receive answer summaries much more often than navigational phrasings [4]. That may sound minor, but for a brand the consequences are enormous. A company may be highly visible in the mode of a direct brand-name query and almost disappear in the mode of a category question, where the decision is made earlier and without any explicit intention to visit the brand’s website. In practice, that means different systems not only answer the same question differently; they also decide differently whether the question deserves a generative answer in the first place. The fifth reason is that systems differ in their criteria for trusting a source. Search Arena shows that users more often prefer answers with a larger number of citations, and that the type of cited sources also influences those preferences [5]. SourceBench emphasizes that source quality directly determines answer reliability [6]. But the question of which sources should count as “quality sources” is resolved differently by each system. For one, large reference hubs matter most; for another, technology and public-discourse platforms; for a third, official documents or commercial catalogs. That is why a brand may win in one environment thanks to strong documentation and lose in another, where the decisive layer is independent reviews. ## Why a single snapshot is almost useless The practical effect of these differences is easy to see in everyday work. Suppose a company sells a complex analytics service for e-commerce. In one answer interface, it may be presented as “a solution for mid-sized and large stores” — because the system relied on the official website, an industry review, and several long-form comparison articles. In another interface, that same brand may look like “an expensive enterprise product” — because the model pulled in a set of external publications about large-scale deployments and ignored the small-business segment. In a third answer, it may disappear altogether, giving way to simpler services if the user’s question was phrased as “what can I start with quickly, without a long implementation.” In all three cases, we are not dealing with falsehood in the strict sense. We are dealing with different modes of selection, emphasis, and generalization. From this follows a very important methodological conclusion: a single snapshot of visibility is almost useless. If a brand checks itself once, in one system, with one query, and in one language, it has not measured the market — it has measured an accident. To understand the real state of affairs, you have to evaluate not only the average result, but also the spread. How many different versions of the brand arise across different systems? How consistently do the key properties recur? How does the citation set change when the wording changes? Does the brand appear in category answers without its name being mentioned directly? Those are the questions that actually reveal a company’s position in the answer environment. For the future Ansmeter database, an almost natural observation scheme suggests itself here. For every query under study, it is worth recording not only the fact of the answer, but also the system, the answer mode, the date, the language, the intent type, the set of citations, the dominant tone, the brand’s place within the composition of the answer, and the number of alternatives that were automatically mixed into the comparison. At that point, the “answer bubble” will stop being a metaphor and become a measurable quantity: it will become possible to see how resilient a brand is to a change of intermediary, and exactly where the divergence begins. ## How to build cross-system observation There is also a deeper business conclusion here. If different systems construct different versions of a brand, then the company’s strategic task is not to achieve absolute uniformity — which is unattainable in principle — but to reduce chaotic variation and increase the share of desirable interpretations. That is achieved not through magical tricks of “optimization for AI,” but through knowledge discipline: consistent wording across owned resources, strong external validation, a clear machine-readable data layer, precise product categorization, and close attention to the kinds of questions in which the brand disappears today. In a certain sense, the “answer bubble” is a new form of market fragmentation. Companies used to fight for a place in search results. Now they also fight for the stability of their entity as it moves from one answer machine to another. That is why a mature brand in 2026 should ask not simply, “what does AI say about us?” but rather, “what versions of us exist across different answer worlds — and which one wins more often than the others?” Only after that question does genuinely modern visibility work begin. ## What seems well established It is well supported that different systems differ in their search infrastructure, source preferences, interface decisions, and synthesis style. That is why the same brand receives different machine-generated versions. ## What remains uncertain or platform-dependent The exact contribution of each mechanism — parametric memory, retrieval, display policy, interface — to the divergence of a specific answer usually remains hidden from external observation. ## Practical implications for brand work The direct rule that follows is simple: checking one system with one wording tells you almost nothing about a brand’s real position. What you need is a series of runs, languages, and platforms. ## Sources - [1] Huang M. et al. Answer Bubbles: Information Exposure in AI-Mediated Search. 2026. https://arxiv.org/abs/2603.16138 - [2] Google Search Central. AI Features and Your Website. 2025-2026. https://developers.google.com/search/docs/appearance/ai-features - [3] Chen M. et al. Navigating the Shift: A Comparative Analysis of Web Search and Generative AI Response Generation. 2026. https://arxiv.org/abs/2601.16858 - [4] Ovadya A. et al. The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale. 2026. https://arxiv.org/abs/2602.13415 - [5] Search Arena: Analyzing Search-Augmented LLMs. 2025. https://arxiv.org/abs/2506.05334 - [6] Zhang Y. et al. SourceBench: Can AI Answers Reference Quality Web Sources? 2026. https://arxiv.org/abs/2602.16942 --- Source: https://ansmeter.com/knowledge-base/update-lag # Article 9. Update lag: how quickly AI systems change their representation of a company after news, a product launch, or a price change *Research article.* **Research question.** What stages make up the lag in a machine answer, and how can one measure the speed at which a brand updates after the underlying facts change. **Evidence type.** Documents from Google Search Central, Microsoft IndexNow and Bing Webmaster Tools, and OpenAI documentation on search crawlers and commerce scenarios. **Data freshness.** This article is based on materials current as of March 2026. > The internet is not instantaneous for AI. Between a change in fact and a durably updated answer lie publication, crawling, indexing, and synthesis; that is why update lag is becoming a new operational discipline of the brand. ## The internet is not instantaneous for an answer system The most treacherous mistake in discussions of brand visibility in AI is to assume that the internet is instantaneous. From a human point of view, that intuition is understandable: the news has been published, the price on the site has changed, the product card has been updated, the press release has gone out to the mailing list. It seems that after that, the world should already “know” the company’s new version. But answer systems operate on their own time. They do not live in the brand’s time, but in the time of crawling, indexing, repeated retrieval, synchronization of data feeds, and, finally, renewed answer synthesis. That is why there is almost always a lag — a delay — between an event and its full reflection in the brand’s machine representation. Sometimes it is measured in hours. Sometimes in days. Sometimes in weeks. And in some cases, longer still. To understand the nature of that delay, it helps to break it down into several layers. The first layer is publication lag: the moment when the company itself actually introduced the change into the canonical source. Very often, a business says, “we’ve already updated the information,” when in fact the update was made only on a single page, without synchronization across documentation, pricing, product cards, and external profiles. The second layer is discovery lag: a search crawler or another technical agent has to notice that the page has changed. The third layer is indexing lag: the change has to enter the index or the platform’s machine-readable infrastructure. The fourth layer is answer-assembly lag: even indexed information does not necessarily surface immediately in a specific AI answer. The fifth layer is source-alignment lag: if external sources continue to describe the brand in the old way, the system may continue, for some time, to hold on to the previous version of the entity as it tries to reconcile conflicting evidence. In its simplest form, the total delay can be written as: L_total = L_pub + L_disc + L_index + L_synth where L_total is the full delay between a change in the company and the appearance of a durably updated machine answer, L_pub is publication lag, L_disc is discovery lag, L_index is indexing lag, and L_synth is synthesis lag. The formula obviously simplifies reality, because some processes may run in parallel. But it is useful precisely as a thinking tool: the brand stops perceiving “updating in AI” as a single magical act and begins to see a sequence of distinct technical and content transitions. ## What Google, Bing, and OpenAI say Official documents from the major platforms confirm this multilayered structure. Google explains that for a page to appear in AI Overviews and AI Mode, it must be indexed and generally eligible to appear in ordinary search with a snippet; there are no special “AI requirements” for that [1]. In other words, before a page can become part of the answer environment, it must pass through the ordinary discipline of search accessibility. Moreover, in the same guidance, Google reminds site owners that indexing and display are not guaranteed even if all requirements are met [1]. In practice, that means that after updating a page, a brand cannot simply assume the work is finished; it still has to wait for the change to be discovered and to become genuinely available to answer modes. The situation is especially visible in e-commerce. Google explicitly recommends combining structured data on the site with a product data feed in Merchant Center, because structured data improves the accuracy with which price, discounts, shipping, and availability are understood, while the product feed provides greater control over update timing, especially for large and frequently changing catalogs [2]. The same guidance states directly that frequent changes in price and availability are exactly what make the data feed especially important [2]. That is a revealing detail. In the classical editorial logic, the site seemed sufficient. In reality, commercial visibility increasingly depends on how quickly and reliably the system receives a machine-readable signal that something has changed. Microsoft and the IndexNow ecosystem make that dependency even more explicit. The official IndexNow site describes the protocol as a way to instantly notify participating search engines about content changes, whereas without such a signal discovery may take anywhere from several days to several weeks [3]. Bing states outright that generative search, real-time shopping, price promotions, restocking, and new product launches raise the requirements for data update speed; fragmented feeds and slow indexing cease to be a minor technical nuisance and become a direct cause of lost visibility [4][5]. This means lag has stopped being merely an inconvenience. It has become a competitive factor. At OpenAI, the logic is the same, though it is expressed through the company’s own infrastructure. Documentation for the OAI-SearchBot crawler says that after a change to robots.txt, the system needs about a day to reconfigure site access for search purposes [6]. That is a small but very important marker: even a simple change in access rules does not take effect instantly. In the commercial layer, OpenAI goes further still and offers a direct product feed specifically intended to allow ChatGPT to “accurately index and display” products with current price and availability [7]. In OpenAI’s shopping help article, the company additionally warns that after a price or shipping change, there may be a delay before the new information is reflected, which is precisely why merchants are offered a direct feed [8]. And in the March 2026 release notes, the company separately reports that it improved product data coverage, freshness, and speed through the Agentic Commerce Protocol [9]. In other words, the leading platforms are not concealing the lag problem — they are building entire product solutions around it. ## What determines the duration of the lag Several strategic conclusions follow from this for brands. First, update lag depends on the type of fact. A change in a company name, core positioning, or product composition is one type of update. A change in price, availability, or return conditions is another. News about a partnership, a funding round, or the release of a research study is a third. These facts have different levels of “machine sensitivity.” Commercial data can usually be accelerated more effectively through feeds, markup, and notification protocols. Reputational and meaning-level changes update more slowly, because they require not only the site to be crawled, but the entire network of external evidence to be reworked. Second, the lag almost always grows when a brand maintains several weakly synchronized sources of truth. For example, the price has already been updated in the catalog, but remains old in the structured data. Availability has been corrected on the site, but Merchant Center has not yet caught up. A new plan has been published in the blog, but has not been added to the comparison table or reflected in the FAQ section. In that state, the system receives not an update, but a conflict. And when faced with conflict, answer systems tend either to become cautious or to rely on the source that appears more reliable and more formalized within their infrastructure. Third, lag cannot be reduced to a single site. Even if a brand updates its own pages very quickly, the external contour may continue to live in the old version for a long time. An analytical article, an industry ranking, a directory, an aggregator, an old comparison with competitors — all of these continue to exist and participate in answer assembly. That is why, in sensitive cases, update work must include not only internal publication, but also a program of external synchronization: updating profiles, catalogs, the press kit, company listings, and sometimes proactively correcting widespread errors on third-party platforms. ## How to measure and reduce lag For Ansmeter’s own research database, update lag is one of the most fertile topics. It can be measured almost in laboratory fashion. It is enough to choose a fact type — for example, a price change, the launch of a new feature, or the release of a major study — record the exact moment when the update appears in the canonical source, and then check at regular intervals how long it takes different AI systems to begin reproducing the new version consistently. Such a design yields not only interesting content, but extremely practical knowledge: which platforms react faster to which types of change, where data feeds work better, where external citations matter more, and where crawling signals are critical. As a result, update lag turns out not to be a minor technical detail, but the heart of a brand’s new operational discipline. In the classic internet, one could afford a certain slowness: the user still came to the site and saw a fresh page. In the answer environment, that is no longer always true. The user encounters the synthesis first. And if that synthesis is assembled from old data, the brand enters the conversation with the market wearing an outdated mask. That is why modern work on AI visibility begins not only with content, but with the speed at which knowledge is updated. Those who know how to shorten the lag gain not simply a fresher site, but a more up-to-date version of themselves in the market’s machine perception. ## What seems well established It is well established that answer updating goes through several stages and may lag in different ways for different types of facts: names, prices, assortment, editorial evaluation, or reviews. ## What remains uncertain or platform-dependent A single uniform update speed for specific platforms and verticals is established much less clearly. These timelines depend on crawl frequency, data availability, query type, and on whether the update also appears in external sources. ## Practical implications for brand work The practical purpose of this article is to move the conversation about freshness away from the level of “we think the system is outdated” and into a measurable log of delays by fact type and by update channel. ## Sources - [1] Google Search Central. AI Features and Your Website. 2025-2026. https://developers.google.com/search/docs/appearance/ai-features - [2] Google Search Central. Share your product data with Google. 2025-2026. https://developers.google.com/search/docs/specialty/ecommerce/share-your-product-data-with-google - [3] IndexNow. How it works. 2025-2026. https://www.indexnow.org/ - [4] Microsoft Bing Webmaster Blog. Keeping Content Discoverable with Sitemaps in AI Powered Search. 2025. https://blogs.bing.com/webmaster/June-2025/Keeping-Content-Discoverable-with-Sitemaps-in-AI-Powered-Search - [5] Microsoft Bing Webmaster Blog. IndexNow Enables Faster and More Reliable Updates for Shopping and Ads. 2025. https://blogs.bing.com/webmaster/June-2025/IndexNow-Enables-Faster-and-More-Reliable-Updates-for-Shopping-and-Ads - [6] OpenAI Developers. Overview of OpenAI Crawlers. 2026. https://developers.openai.com/api/docs/bots/ - [7] OpenAI Developers. Agentic Commerce - Products. 2026. https://developers.openai.com/commerce/specs/file-upload/products - [8] OpenAI Help Center. Shopping with ChatGPT Search. 2026. https://help.openai.com/en/articles/11128490-shopping-with-chatgpt-search - [9] OpenAI Help Center. ChatGPT Release Notes - March 24, 2026 Shopping updates. 2026. https://help.openai.com/en/articles/6825453-chatgpt-release-notes --- Source: https://ansmeter.com/knowledge-base/access-economics # Article 10. Access economics: crawling, indexing, training, and the brand’s right to manage its presence *Research article.* **Research question.** How should one distinguish among content access modes — search, AI answer, training, and agentic use — and why is this now an economic question, not just a technical one. **Evidence type.** Google and OpenAI documents on crawlers and access rights, Cloudflare materials, and research on the changing economics of content consumption. **Data freshness.** The facts and examples reflect the market regime of 2025–2026. > In the new environment, granting a bot access is no longer an unambiguous good. Content increasingly functions as an asset with different access modes, and the brand has to distinguish among indexing, answer use, and training. ## The old contract between the site and the bot has broken down In the old web economy, allowing a bot onto a site was treated as an almost unconditional benefit. Search crawling led to indexing, indexing led to visibility, visibility led to traffic, and traffic led to advertising, subscriptions, or sales. It was a crude model, but it worked long enough to become almost a natural law of the internet. Answer systems disrupted precisely that law. Now the same text can participate in several chains at once: it can support a search answer, serve as training material for a model, be used to “ground” an answer at query time, or be retrieved through a direct user-initiated action. These chains look similar technically, but they differ economically. Which means the question of access to content stops being binary. It no longer sounds like “do we let the bot in or not?” It breaks down into a harder question: “which bot, for what purpose, and on what terms are we prepared to admit?” To discuss this seriously, one has to distinguish at least four access modes. The first is crawling and indexing for search visibility. The second is the use of content to train future models. The third is the use of a search index or a web document to answer at the moment of the query — in other words, to ground the answer operationally. The fourth is user-initiated access to the site, when the system itself acts as an intermediary for the user’s request. If these modes are mixed into one mass, the brand loses control and starts making decisions based either on vague fears or, conversely, on naive optimism. ## Four access modes and their new separation Google and OpenAI have already, in effect, formalized this distinction in their own rules. Google Search Central states directly that AI search features — AI Overviews and AI Mode — are governed by the same access rules as ordinary search: the key agent here remains Googlebot, while visibility restrictions in search AI features rely on familiar mechanisms such as `nosnippet`, `data-nosnippet`, `max-snippet`, or `noindex` [1]. At the same time, Google emphasizes that `Google-Extended` is a separate token through which a publisher can control the use of content for training future generations of Gemini and for grounding in Gemini Apps and certain cloud scenarios; `Google-Extended` does not affect inclusion in Google Search and is not a ranking signal [2]. A very important conclusion follows from this: at Google, search visibility and model training have already been institutionally separated. It is no longer intellectually honest to say simply “we allowed Google” or “we blocked Google” without specifying which process is actually meant. OpenAI frames a similar distinction even more explicitly. OpenAI’s documentation says that OAI-SearchBot is responsible for the appearance of sites in ChatGPT’s search functions, GPTBot is used for training foundation models, and ChatGPT-User handles actions initiated by the user [3]. More than that, OpenAI explicitly says that a webmaster can allow OAI-SearchBot so that the site can participate in search answers while blocking GPTBot so that the content is not used for training [3]. In essence, this creates a new publisher right: the right to distinguish between useful visibility and unwanted value extraction. This is exactly the basis on which the new access economics emerges. In 2025, Cloudflare put the problem in especially stark terms: old search crawlers and publishers were linked by a symbiotic exchange, whereas many new training bots consume content while returning almost no traffic [4]. According to Cloudflare, in June 2025 Google crawled sites roughly 14 times for every one referral, whereas OpenAI’s crawl-to-return ratio stood at 1,700:1 and Anthropic’s at 73,000:1 [4]. Even if one allows for the fact that some referrals from apps may not be captured in the `Referer` header, the asymmetry is too large to dismiss as statistical noise [4]. It means that the old informal contract — “you get content, we get audience” — no longer operates automatically in many AI scenarios. ## From total blocking to differentiated governance But this is also where the brand risks falling into the opposite extreme: the temptation of total blocking. Such a decision may look morally clear, yet it is not always economically sound. If all forms of access are blocked, the result may be not only exclusion from training, but also the loss of some channels of visibility, research, and sales. There are already early empirical signals that blocking bots may be associated with lower traffic for major publishers relative to those that do not block access, although such results still require cautious interpretation [5]. The point is not that blocking is forbidden. The point is that blocking has ceased to be a neutral defensive gesture. It has become a strategic choice with multiple consequence paths. That is why a mature brand position has to be differentiated. If a company wants to be visible in ChatGPT Search but does not want its texts used to train future models, that is already technically possible through separate rules for OAI-SearchBot and GPTBot [3]. If a brand has no objection to participating in Google Search and AI Overviews but does not want content to be used for Gemini training, that can be expressed through a combination of allowing Googlebot and restricting Google-Extended [1][2]. In other words, the market is gradually moving toward a regime of fine-grained tuning of access rights rather than a crude “yes” or “no.” Against that backdrop, attempts to turn access to content into a transactional object are of particular interest. In the summer of 2025, Cloudflare introduced a pay per crawl model in which a domain owner can choose one of three modes for a specific bot: allow access for free, charge for crawling, or block completely [6]. For now, this is more of an infrastructure experiment than a mass standard. But its significance is hard to overstate. For the first time, it makes visible the fact that crawling no longer has to remain a free gift. If an AI company extracts value from someone else’s content outside the logic of traffic return, then the question of the price of that access becomes entirely rational. There is another practical side to the problem that is rarely discussed in public. Many sites still do a poor job of formalizing their own rules of engagement with bots. Cloudflare notes that only about 37% of the largest domains have a robots.txt file at all, and among existing robots.txt files, restrictions on the key AI agents are surprisingly rare [4]. That means a significant share of the internet entered the new era without having articulated its own legal and technical position. Companies debate AI as a global cultural problem, yet at the infrastructure level they have not even stated their own “yes” or “no” in a format machines can read. ## Content as an asset with access terms For brands, this is not an abstract legal issue. It is a question of the cost and role of content. Some materials are created as marketing assets for maximum distribution. Others are research assets that required investment, so the brand may want to limit free extraction. Still others function as commercial catalogs, where up-to-date visibility matters most. And others serve as operational documentation that should be shown only in specific scenarios. A modern access strategy has to distinguish among at least these classes and assign them different participation modes in the answer environment. For Ansmeter, the topic of access economics is especially rich from a research perspective. It allows one to build an observation base across several layers at once: which agents actually access the site, how robots.txt is configured, where access is allowed and where it is restricted, how that affects the brand’s visibility in answer systems, and how crawl volumes relate to actual return traffic or commercial interest. Over time, this material may become one of the most valuable assets in the entire base, because much of the market still discusses AI access in moral categories rather than in terms of a measurable architecture of value exchange. The main conclusion here is fairly strict. In the new environment, content is no longer merely a message. It is an asset with several channels of value extraction. It can bring in a customer, shape a machine answer, train a future model, or become a good for which the publisher will sooner or later ask compensation. That is why the brand’s right to manage its presence is not the right to disappear. It is the right to choose the specific mode in which its knowledge will participate in the economics of AI. And in the coming years, the winners will not be those who are the loudest in outrage or enthusiasm, but those who build a calm, precise, and technically competent access policy for their own knowledge. ## What seems well established It is already well established that major platforms separate search crawling from training, and that a brand can configure access to those modes differently. The economic asymmetry between crawling and returned traffic has also been publicly documented. ## What remains uncertain or platform-dependent What is far less certain is what market mechanisms for charging for crawling will ultimately become and how quickly they will turn into a mass norm. Here the market is still in an experimental stage. ## Practical implications for brand work For a company, the practical implication is that access policy needs to become part of content strategy and engineering architecture, not a random set of lines in robots.txt. ## Sources - [1] Google Search Central. AI Features and Your Website. 2025-2026. https://developers.google.com/search/docs/appearance/ai-features - [2] Google for Developers. Google's common crawlers - Google-Extended. 2025-2026. https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers - [3] OpenAI Developers. Overview of OpenAI Crawlers. 2026. https://developers.openai.com/api/docs/bots/ - [4] Cloudflare Blog. Control content use for AI training with Cloudflare’s managed robots.txt and blocking for monetized content. 2025. https://blog.cloudflare.com/control-content-use-for-ai-training/ - [5] Zhao H., Berman R. The Impact of LLMs on Online News Consumption and Production. 2026. https://arxiv.org/abs/2512.24968 - [6] Cloudflare Blog. Introducing pay per crawl: Enabling content owners to charge AI crawlers for access. 2025. https://blog.cloudflare.com/introducing-pay-per-crawl/ --- Source: https://ansmeter.com/knowledge-base/machine-readable-commerce-stack # Article 11. Machine-readable commercial infrastructure: markup, product data feeds, and catalogs as a language AI can understand *Research article.* **Research question.** Why the catalog, markup, and product data feed are no longer a secondary technical add-on, but the language through which the brand speaks to AI about its commercial reality. **Evidence type.** Google Search Central documentation on ecommerce and structured data, and OpenAI documentation on product feeds and shopping scenarios. **Data freshness.** The factual basis of the article was updated as of March 2026. > In a commercial environment, text without structured data becomes too poor a language. An answer system needs machine-readable signals about product identity, price, availability, return policy, variants, and seller context. ## The editorial and machine sides of the catalog Many companies still talk about visibility in AI as if it were a purely editorial task. The assumption is that they simply need to write better copy, articulate their advantages more clearly, and present the product page more neatly. All of that does matter. But for commercial presence in the answer environment, good writing alone is no longer enough. Answer systems read not only paragraphs; they read structures. They care about price, availability, shipping, returns, product variants, seller identity, the relationship between the organization and the offer, and the degree to which those data are current. And the more of the user’s decisions happen before visiting the site, the more important machine-readable commercial infrastructure becomes—that layer of formalized information a machine can interpret without prolonged guesswork. Google says this with almost no diplomacy. In its ecommerce documentation, the company recommends combining structured data on the site with a product data feed in Merchant Center rather than relying on only one of those channels [1]. Structured data improves the accuracy of how price, discounts, shipping, and availability are understood, while Merchant Center gives the company more control over assortment and update timing, especially if the catalog is large and changes often [1]. In effect, Google is acknowledging that human-written text and machine-readable product discipline are now two different, but equally important, layers of the brand’s commercial entity. It is important here to avoid a common misconception. Google says two things at once that, at first glance, seem contradictory: there is no special “AI markup” for AI Overviews and AI Mode, yet structured data, textual accessibility of key content, and Merchant Center freshness still matter critically [9]. In reality, there is no contradiction. Platforms are not waiting for some magical new tag “for AI”; they are waiting for commercial information to be described in the formal structures that search and shopping systems already know how to read. In other words, machine-readable infrastructure is powerful precisely because it is not a hack. It is a discipline of clarity, not a trick for getting around the rules. For an AI answer, this is critical. Human-written text is good at explaining meaning: what makes a product different, who it was created for, and what jobs it solves. But when a user asks, “how much does it cost,” “is it in stock,” “what are the return terms,” “when will it arrive,” or “which version is right for a particular scenario,” the system should not have to pull that information from free-form description every time like an archaeologist working through loose soil. It needs reliably marked anchor fields. That is exactly why Google’s documentation separately introduces schemas for shipping terms, return policy, product variants, and the store’s organizational context [2][3][4][5]. These elements may look dry at first glance, but they are in fact the language in which the brand speaks to the machine about commercial reality. ## What makes a commercial offer machine-readable From a practical standpoint, commercial machine readability consists of six blocks. The first block is seller identity. Who exactly is selling the product or service? How are the brand, the legal entity, the site, organization profiles, and the catalog connected? The second block is the identity of the product itself: name, model, variant, variant group, attributes, and category membership. The third block is the offer itself: price, currency, availability, discount, and item condition. The fourth is logistics: shipping, delivery timing, geography, and inventory status. The fifth is post-purchase policy: returns, exchanges, and warranty. The sixth is freshness: when all of this was last updated and through which source contour the platform received the signal that something had changed. If even two or three of these blocks are described vaguely, the system is forced to reconstruct the missing pieces from indirect clues. And where the machine has to guess, the brand loses control. In recent years, Google has been developing exactly this view in a very consistent way. The company separately recommends an organization-level description for return policies so that the same long constructions do not have to be repeated on every product card, while also increasing the chances that the information will appear in brand profiles and knowledge panels [4]. Support for product variant groups is likewise framed as a special class of entity that makes it possible to display variants of the same model more accurately [5]. This leads to an important conclusion: machine-readable infrastructure is needed not only for individual product cards, but also for the brand’s overall commercial self-description. In 2026, OpenAI effectively arrives at the same conclusions, only within its own shopping scenarios. In the Agentic Commerce documentation, the company offers a structured product feed so that ChatGPT can “accurately index and display” products with current price and availability [6]. The shopping documentation states that merchant ranking depends on metadata about the product and the seller—including availability, price, quality, whether the seller is the manufacturer or the primary seller, and whether Instant Checkout is enabled [7]. What is especially interesting here is that the commercial logic of the answer stops being purely textual. A brand enters the new selection not only because it is well described in prose, but because its data are fit for machine comparison. This radically changes the meaning of the familiar work around a catalog. In the past, the catalog was often treated as a technical appendage to “real” marketing. Now it becomes part of the brand’s argumentation. If the catalog lacks a clear structure for product variants, AI may assemble the product line incorrectly. If shipping and return terms are not presented anywhere in machine-readable form, the system will understand purchase risk less well. If price and availability are updated slowly, the brand loses not in text, but in the speed of truth. And if product data are stored in several unsynchronized places, what the machine receives is not a brand, but a contradiction. In this context, it is especially instructive that both Google and OpenAI emphasize the synchronization problem. Google warns that using both the site and Merchant Center at the same time can lead to conflicts and lags, so it is useful to enable automatic item updates, especially when price or availability changes frequently [1]. OpenAI writes directly that there may be a delay when prices and shipping terms are updated, and that a direct feed is needed precisely to reduce the gap between the catalog’s actual state and what ChatGPT shows [7]. In other words, the central task here is not simply to “mark up the site,” but to build a coordinated data flow between the brand’s internal systems and external answer platforms. ## Synchronization as a hidden discipline For Ansmeter, this topic is especially valuable because it connects the marketing conversation to the company’s engineering reality. Visibility in AI stops being a matter only for editors and organic traffic specialists. It now includes product information systems, catalog update processes, data quality control, feed synchronization, card architecture, and the discipline of the corporate entity registry. In management language, that means one thing: commercial visibility in the answer environment becomes a function not only of content, but of the business’s operational maturity. From a research perspective, this opens an almost boundless field. One can compare brands within the same category by the completeness of their machine-readable description, examine which elements most often correlate with correct display in AI scenarios, measure how the presence of structured data and feeds affects the accuracy of pricing and logistics answers, and track which types of errors arise most often when an organization-level return policy or product groups are missing. Such a body of data would be useful not only as a media asset, but also as the foundation for a consulting product. ## The engineering contour as a condition of marketing But the main conclusion lies deeper than the technical layer. Machine-readable commercial infrastructure is not a way to please the algorithm. It is a way to make the brand’s business entity clear enough that the machine does not distort it on the way to the user. The market is gradually moving away from a world in which the user went to the site and manually assembled a picture from scattered tabs. In the new world, part of that assembly is performed by an intermediary. If the brand has not given that intermediary a precise language for describing its offer, the intermediary will inevitably start filling in the gaps. And where that gap-filling begins, controllability ends. That is why the best commercial brand of the next few years will win not only through product quality and sharp positioning, but through the engineering of its own truth. Whoever can express assortment, price, terms, and the seller’s role in machine-readable form gives AI less room for error and more grounds for confident display. In the answer environment, this is no longer a secondary technical detail. It is a new grammar of trust. ## What seems well established It is well established that both Google and OpenAI are moving toward a more structured way of representing commercial data. Without machine-readable infrastructure, a brand risks being understood only partially or in an outdated way. ## What remains uncertain or platform-dependent Less certain is which combinations of fields and structures will prove most advantageous in each vertical, and how quickly markets will develop stable best practices for this new commercial layer. ## Practical implications for brand work The practical value here is that the catalog needs to be designed as part of the brand’s semantic infrastructure. It is no longer a secondary technical appendage, but a condition for an accurate commercial answer. ## Sources - [1] Google Search Central. Share your product data with Google. 2025-2026. https://developers.google.com/search/docs/specialty/ecommerce/share-your-product-data-with-google - [2] Google Search Central. Structured data for shipping details. 2025-2026. https://developers.google.com/search/docs/appearance/structured-data/shipping-details - [3] Google Search Central. Structured data for return policies. 2025-2026. https://developers.google.com/search/docs/appearance/structured-data/return-policy - [4] Google Search Central Blog. Supporting organization-level return policies and loyalty programs. 2025. https://developers.google.com/search/blog/2025/05/organization-level-return-policies-and-loyalty-program - [5] Google Search Central Blog. New merchant listing support for product variants. 2024-2025. https://developers.google.com/search/blog/2024/02/product-variants - [6] OpenAI Developers. Agentic Commerce - Products. 2026. https://developers.openai.com/commerce/specs/file-upload/products - [7] OpenAI Help Center. Shopping with ChatGPT Search. 2026. https://help.openai.com/en/articles/11128490-shopping-with-chatgpt-search - [8] OpenAI Help Center. ChatGPT Release Notes - March 24, 2026 Shopping updates. 2026. https://help.openai.com/en/articles/6825453-chatgpt-release-notes --- Source: https://ansmeter.com/knowledge-base/external-authority-vs-own-site # Article 12. External authority versus the brand’s own website: which sources actually shape a brand’s right to be recommended *Research article.* **Research question.** Which external sources, exactly, give a brand the right to be recommended, and why the company’s own website becomes insufficient without an external contour of validation. **Evidence type.** Research on the quality of cited sources, comparative work on AI search, and market observations about new customer paths. **Data freshness.** The factual basis was compiled from research and official materials from 2025–2026. > The website remains the brand’s anchor document, but the right to be recommended is increasingly shaped elsewhere — where independent sources validate the company’s category, properties, and reliability. ## The website is no longer the only witness Any brand is naturally inclined to treat its own website as the main place of truth about itself. And the logic is understandable: it is there that the company can give the precise product name, articulate its positioning, describe features, pricing, constraints, and case studies. But in the answer environment, the canonical source ceases to be the sole judge of its own correctness. The website matters, but it no longer holds a monopoly on trust. For a system to recommend a brand, it is not enough for the system to hear how the brand talks about itself. It needs to see how that knowledge is validated, clarified, repeated, or contested from outside. The best way to describe this shift is not as a decline in the role of the website, but as a change in its function. The website remains the primary document of identity, but the right to be recommended is increasingly distributed across several classes of sources. McKinsey writes in its research on the new age of AI search that brand-owned websites often account for only 5–10% of the sources on which answer systems rely; the rest comes from a broad mix of editorial, partner, user, and other external materials [1]. That figure is not a universal law in itself, but it captures the direction of the shift well. In the answer environment, a brand has to be not only self-described, but externally validated as an entity. Why do external sources carry such weight? Above all because they perform different epistemic functions. A company’s own website is good at establishing the official version: who we are, what we sell, how it works. But it is poorly suited to creating independent trust in its own strengths. If a brand calls itself “a leader,” “the most accurate,” or “the best solution for the enterprise segment,” an answer system is not obliged to treat that as an established fact. For such a claim to become part of public machine knowledge, it needs external carriers — research, reviews, rankings, publications, case studies on independent platforms, professional communities, and sometimes government or academic documents. ## What the research literature says about source selection There is, however, an important caveat here. Answer Bubbles shows that generative answers gravitate disproportionately toward certain document types — for example, Wikipedia and longer texts — while some social and negatively toned sources end up underrepresented [4]. So it is not enough for a brand simply to “be mentioned somewhere outside.” It matters which external validations are actually more likely to enter the field of vision of answer systems, and which ones remain out on the periphery of machine attention. In the new environment, the distribution of authority becomes not only reputational, but also interface-driven. Current research confirms that source selection in answer systems is not random and has a direct effect on trust. In SourceBench, the authors emphasize that the quality of web sources directly determines answer reliability, while users tend to trust answers with links even if they do not go on to verify those links themselves [2]. Search Arena adds an important detail: users prefer answers with a larger number of citations, and the type of source also influences preference — links to technology, public-interest, and discussion platforms are often perceived more favorably than overloaded or overly general reference sources [3]. A subtle but important conclusion follows from this: a brand’s right to be recommended is built not from one “best” source, but from a configuration of validations that the system considers sufficient for a convincing synthesis. That does not mean the brand’s own website becomes secondary. On the contrary, without it, external validations often lose their anchor. In the answer environment, the website performs at least three irreplaceable functions. The first is canonization: it fixes official names, categories, characteristics, and relationships between entities. The second is detail: it provides a depth that external sources rarely offer in full. The third is alignment: it serves as the place where discrepancies among different external versions of the brand can be checked. But all three functions work fully only when the external source contour does not contradict the website too sharply and does not leave the system in a source vacuum. ## Classes of external authority and their unequal strength The problem is that many companies spent decades operating within a logic where the external contour was treated as optional. The main thing was a good website, and everything else was simply a pleasant bonus. In classic search, that stance could still produce results, especially if the brand already had demand power. In the answer environment, it is not enough. Answer Bubbles shows that different systems have pronounced biases in source selection; some document types are systematically overvalued, while others are underrepresented [4]. The paper Navigating the Shift further shows that generative answers diverge noticeably from traditional search in the types of domains they use, the freshness of the information, and the balance between owned and external sources [5]. For a brand, that means it can no longer assume that the website will automatically become the center of all machine reasoning about the company. It is useful here to distinguish among several classes of external authority. The first class is institutional: government domains, regulatory materials, academic publications, and professional standards. These rarely create emotional attractiveness around a brand, but they work well for factual reliability and for belonging to a serious category. The second class is editorial: industry media, reviews, rankings, interviews, and analysis. These often shape the external interpretation of the brand’s role in the market. The third is community-based: forums, Q&A spaces, expert communities, and user discussions. This layer is noisier, but it is exactly what helps the machine understand the language of real demand and real usage scenarios. The fourth is commercial-reference: catalogs, company profiles, marketplace listings, supplier databases, and product aggregators. Here precision, consistency, and freshness matter. The brand’s own website should not replace these layers; it should connect them into a coherent, non-contradictory system. What matters especially is that external authority and media noise are not the same thing. A large number of weak mentions on irrelevant platforms does not necessarily help a brand. On the contrary, answer systems may prefer a handful of strong, substantive, and independent validations to dozens of superficial publications written from the same template. Moreover, The Rise of AI Search shows that AI answers tend on average to surface the largest sites more often and to cite the web’s long tail less often [6]. That makes the question of the quality of the external contour even sharper. If the system is already compressing diversity, the brand has fewer chances to float to the surface by accident. It has to build a more deliberate architecture of authority. In practice, this changes the meaning of content strategy. A strong brand website is no longer the end goal; it becomes the core around which a network of validating documents from different origins must be built. If a brand says it is good for complex enterprise implementations, ideally that should be visible not only on the product page, but also in independent case studies, industry reviews, user discussions with real implementation detail, comparative materials, and, where possible, in business profiles with a clear description of the customer segment. If a brand wants to be associated with a certain category, that category has to be anchored not only in its own headlines, but also in the market’s external language. ## How to build an authoritative external contour For Ansmeter, this topic is especially fertile because it can be researched both broadly and deeply. Broadly — by comparing which types of external sources appear most often in answers across categories. Deeply — by analyzing which specific combinations of sources give a brand not just mention, but the right to be recommended. For example, what works better for B2B: an editorial review plus a case study plus a company profile, or an official website plus a forum plus a catalog? Which type of external validation most often secures a brand’s role near the top of a comparative answer? Questions like these move the conversation about “reputation on the internet” out of metaphysics and into an applied dimension. The final conclusion is not especially pleasant for brand-centric consciousness, but it is realistic. In the answer environment, the website says: “this is who we are.” External authority replies: “we confirm it” — or does not. The answer system, in turn, builds a recommendation where enough alignment arises between those two voices. That is why a brand’s right to be recommended is born not in isolation, but in a network. A strong website remains necessary. But the winners are those who have learned to turn self-description into a publicly validated and machine-stable reality. ## What seems well established It is well established that in answer systems, an external source often serves as a validating and legitimizing layer rather than merely an additional mention. The quality and type of those sources affect both trust and the wording of the answer. ## What remains uncertain or platform-dependent What is much less clearly defined is any universal ranking of all source types. The real weight of institutional, editorial, and user validation depends on the topic, the level of risk, and the architecture of the system. ## Practical implications for brand work The practical point of the article is that content strategy can no longer be purely internal. A brand needs to build not only a strong website, but also a clear external contour of evidence. ## Sources - [1] McKinsey & Company. New front door to the internet: Winning in the age of AI search. 2025. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search - [2] Zhang Y. et al. SourceBench: Can AI Answers Reference Quality Web Sources? 2026. https://arxiv.org/abs/2602.16942 - [3] Search Arena: Analyzing Search-Augmented LLMs. 2025. https://arxiv.org/abs/2506.05334 - [4] Huang M. et al. Answer Bubbles: Information Exposure in AI-Mediated Search. 2026. https://arxiv.org/abs/2603.16138 - [5] Chen M. et al. Navigating the Shift: A Comparative Analysis of Web Search and Generative AI Response Generation. 2026. https://arxiv.org/abs/2601.16858 - [6] Ovadya A. et al. The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale. 2026. https://arxiv.org/abs/2602.13415 --- Source: https://ansmeter.com/knowledge-base/category-drift # Article 13. Category Drift: How a Brand Loses Not Only to a Competitor, but to Someone Else’s Framing of the Choice *Research article.* **Research question.** How can a brand lose not to a direct competitor, but to someone else’s choice frame, when the system changes the name of the task and assembles a different set of alternatives. **Evidence type.** Google documents on AI Mode, research on the impact of query type on AI search, and Microsoft materials on the new logic of conversion and the shorter user path. **Data freshness.** The platform and research evidence discussed here dates from 2025–2026. > In the answer environment, competition begins before brands are compared directly. The system may first rename the user’s task and only then select alternatives—already within a new category that disadvantages the brand. ## A brand can lose before it is ever compared with a competitor When companies think about competition on the internet, they usually imagine direct rivalry between brands. A user asks about one product, a search engine or marketplace shows neighboring products, and the familiar battle for attention begins. In the AI answer environment, that model turns out to be too narrow. Here a brand can lose before it ever meets a competitor—simply because the system frames the question differently. The user asks not about a brand, but about a task; the machine translates the question into the language of a category; the category then breaks down into a set of criteria; and only within those criteria do other players appear, including ones the user did not initially have in mind. This is category drift: defeat not in a comparison between brands, but in the frame within which the system decides who counts as relevant. That mechanism is built into the very nature of modern AI search systems. Google writes that AI Mode is especially useful for complex comparisons and nuanced questions, and that AI Overviews and AI Mode can use query fan-out, breaking a query into subtopics and additional searches across related aspects [1]. In commercial choice terms, this means the following: if a user asks a question about solving a task, the system is under no obligation to limit itself to the brands the user already knows. It can first determine the latent category, then identify the criteria, and only afterward assemble a set of alternatives that fit those criteria. In that logic, a brand can disappear long before any direct comparison with competitors begins. Empirical evidence confirms that query type is decisive here. In The Rise of AI Search, researchers showed that AI answers occur far more often for questions than for navigational queries, where the user already knows where they want to go [2]. In other words, a brand is relatively safe when demand has already formed in its favor: a person types the company name, the product name, or something very close to a direct path to it. But in the early and middle stages of choice—where the user asks “what is best for this task,” “where should I start,” or “which tool fits these constraints”—the mediator has maximum freedom to define the frame. And that is precisely where the most painful loss of attention share occurs. The significance of this mechanism is growing not just in theory, but in mass user behavior. McKinsey writes that roughly half of consumers already use AI-assisted search, and 44% of those users call it their primary source of information when making purchase decisions [5]. If that is the case, then category drift is no longer a rare interface error. It becomes a systematic point at which demand is redistributed. More and more often, the user receives the first frame of choice not from the brand, and not from their own path through search results, but from a mediator that decides what language the task will be described in at all. ## Four forms of category drift Category drift can occur in at least four forms. The first is a change in the name of the task. The brand believes it operates in one category, while the system sees the user need in another. A company may sell, for example, an “intelligent analytics environment for commerce,” while AI translates that into the simpler phrase “a reporting tool for online stores.” The second is a change in criteria. The brand built its positioning around accuracy, depth, and integration, while the machine decides that in this question the decisive criteria are ease of launch and a low barrier to entry. The third is a narrowing or widening of the set of alternatives. The system may unexpectedly mix in services from adjacent categories if, in its view, they answer the user’s request better. The fourth is the displacement of the brand by a description of a solution class. In that case, the answer may omit company names altogether and remain at the level of “it is better to choose tools with such-and-such properties.” Formally, the brand did not lose to a direct competitor. But in practice, it has already been pushed out of the moment of choice. In the answer environment, this is especially dangerous for companies with complex or overly technical self-descriptions. A user rarely comes to an answer system with the brand’s ready-made terminology. They describe the problem in ordinary language: “I want something faster to implement,” “I’m not ready for a heavy integration,” “I need a tool the team can understand without a separate analyst,” “I’m looking for a solution for a midsize business, not a huge corporation.” If a brand has spent years speaking about itself in the language of internal categories, without connecting that language to the real language of demand, the system will easily place it in someone else’s bucket—or fail to see any reason to place it anywhere at all. In its materials on the new metrics of AI search, Microsoft writes that the user path becomes shorter, but at the same time more deeply integrated into the answer environment itself: intent is refined at every turn of the dialogue, and part of the choice happens before the site is ever visited [3]. That detail has direct bearing on category drift. The more of the decision is made inside the conversation, the greater the influence of the criteria and comparisons that the system itself brings to the surface. A user may begin the conversation with a relatively neutral task, but by the second or third turn end up inside a frame where entirely different classes of solutions are being considered. The problem is compounded by the fact that brands often measure themselves through branded demand and draw false conclusions from it. If users who already know the company continue to find it by name, that creates an impression of resilience. But it is no accident that Google introduced a separate filter for branded and unbranded queries in Search Console [4]. In doing so, the platform effectively acknowledges that these are two different worlds with two different growth logics. A branded query shows the strength of knowledge about the company that has already formed. An unbranded query shows the brand’s ability to appear where the user has not yet decided whom they are asking about. In the AI answer environment, it is this second world that becomes the main battlefield. ## Why the shorter user path is more dangerous For Ansmeter, category drift could become an especially strong research series. For each industry, one can assemble a corpus of real user phrasings and see which categories different systems translate them into. How often does the brand remain inside the original category? How often does the system substitute a different set of criteria? What kinds of adjacent solutions get mixed into the comparison? How much changes when one or two phrases in the task wording are altered? Studies of this kind would quickly show that losing in AI rarely looks like “a competitor took our place.” Much more often, it is a quiet loss of the right to be included in the choice frame at all. The practical response to this problem begins with rebuilding the brand’s own language. A brand needs to anchor itself not only in the category it considers correct, but also in adjacent user phrasings of the task. That means content, documentation, external reviews, comparison materials, and machine-readable descriptions need to carry not only the brand’s internal positioning, but also bridges to the real language of demand. If the product is suitable “for a quick start without lengthy implementation,” that should be stated as clearly as its architectural strengths. If the company wants to be associated not only with the large enterprise segment, that needs to be confirmed by external cases and descriptions, rather than remaining an internal brand aspiration. ## How to rebuild brand language and measure drift There is also a broader strategic conclusion. In classic search, a brand could afford to live, at least in part, inside its own category and wait for the user to arrive there on their own. In the answer environment, the mediator does not wait. It builds the bridge from problem to solution itself—and therefore decides which categories count as relevant. Consequently, the modern battle for visibility is a battle not only for mention of the brand, but also for the right to define the vocabulary of the task itself. Whoever loses in the vocabulary begins losing before products are even compared. That is why category drift is one of the most important themes for a mature understanding of the AI market. It shows that a brand’s new kind of defeat can be almost invisible to classical analytics. The site has not lost positions for its own name. A competitor does not appear to have beaten the brand head-on. But demand has already leaked into a different frame, where choices are made according to someone else’s criteria and among someone else’s players. In this new environment, the winner is not only the one who is known, but above all the one who has managed to become the natural answer to the user’s task before the machine renames that task in its own way. ## What seems well established It seems well established that answer systems actively decompose complex questions into sub-tasks and can thereby change the frame of choice. That increases the risk that a brand will be compared with the wrong alternatives—or not named at all. ## What remains uncertain or platform-dependent What is less well understood is how stable this effect is across industries and languages. Across different tasks, locales, and user phrasings, category drift may vary substantially. ## Practical implications for brand work The practical conclusion is that a brand needs to establish itself not only in its own self-description, but also in the language of the user’s tasks. Otherwise, the mediator will define the market frame on the brand’s behalf. ## Sources - [1] Google Search Central. AI Features and Your Website. 2025-2026. https://developers.google.com/search/docs/appearance/ai-features - [2] Ovadya A. et al. The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale. 2026. https://arxiv.org/abs/2602.13415 - [3] Microsoft Bing Webmaster Blog. How AI Search Is Changing the Way Conversions are Measured. 2025. https://blogs.bing.com/webmaster/November-2025/How-AI-Search-Is-Changing%E2%80%AFthe%E2%80%AFWay%E2%80%AFConversions%E2%80%AFare-Measured - [4] Google Search Central Blog. Introducing the branded queries filter in Search Console. 2025-2026. https://developers.google.com/search/blog/2025/11/search-console-branded-filter - [5] McKinsey & Company. New front door to the internet: Winning in the age of AI search. 2025. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search --- Source: https://ansmeter.com/knowledge-base/seo-and-ai-visibility # SEO and AI Visibility: What Carries Over, What Does Not, and Where Familiar Optimization Can Backfire **Research question.** Which skills, practices, and infrastructures from classical SEO genuinely help in the AI answer environment, which ones stop working, and what new requirements emerge. **Evidence type.** The academic GEO study (Princeton, Georgia Tech, IIT Delhi), Google Search Central documents, and empirical data from Seer Interactive, Semrush, Pew Research Center, and Similarweb on CTR, zero-click behavior, and AI Overviews. **Data freshness.** The empirical data discussed here comes from 2025–2026; the platform documents are current as of March 2026. ## The familiar world is not disappearing, but it is shrinking When the conversation turns to AI visibility, the first question a business owner or marketer almost always asks is the same: “We have strong SEO—does that help or not?” The answer is not as simple as one might like. Some skills and infrastructures from classical search genuinely continue to work in the new environment. Some are losing significance. And some habits accumulated over years of optimization are capable not merely of failing to help, but of actively getting in the way. To understand why, it is enough to look at the change in user behavior. According to Similarweb, the share of search queries that end without a single click to an external site rose from 56% in May 2024 to 69% by May 2025 [1]. A Pew Research Center study based on 68,000 real search queries showed that when an AI Overview was present, users clicked on results in 8% of cases, compared with 15% when it was absent [2]. Seer Interactive, using a sample of 25 million organic impressions, found that organic CTR for queries with AI Overview fell to 0.61%, compared with 1.76% without it [3]. This is not noise or statistical coincidence. More and more often, the user gets an answer without leaving the interface of the search system or answer system. Does that mean SEO is dead? No. Google Search Central states explicitly that there are no additional requirements for appearing in AI Overviews and AI Mode—the same fundamentals of search optimization still matter [4]. A page must be indexed and eligible to appear in ordinary search. But “still matter” and “are sufficient” are two different things. SEO fundamentals open the door; they do not guarantee that the brand will make it inside the answer. ## What carries over from SEO and continues to work Technical site accessibility. If a search crawler cannot crawl the pages, they will not enter the index and therefore cannot become a source for an AI summary. A valid robots.txt, an XML sitemap, properly functioning canonical URLs, and fast load times all remain basic requirements for entry. Google emphasizes that for a page to become a supporting link in AI Overviews or AI Mode, it must be indexed and allowed to appear with a text snippet [4]. Content quality and expertise. E-E-A-T (experience, expertise, authoritativeness, trustworthiness)—the set of criteria Google uses to evaluate the usefulness of content—not only has not lost relevance, but has become even more important. Answer systems prefer sources that can be verified and formulations that read like expert judgment rather than marketing copy. In the academic GEO study published by Princeton, Georgia Tech, and IIT Delhi, three of the nine tested optimization strategies performed best: adding specific statistical data, citing authoritative sources, and including expert quotations [5]. This is a direct continuation of what strong SEO has taught for years. Structured data and markup. Schema.org, Open Graph, and machine-readable descriptions of products and services all help the system identify an entity more quickly and more accurately. In the answer environment, this layer becomes not merely a useful addition, but part of the language through which the brand speaks to the machine. Link profile and external authority. External links still signal trust. But—and this is where the differences begin—the key issue is no longer link density so much as which sources those links come from. Independent reviews, industry media, and analytical publications carry more weight than links from directories and aggregators. ## What stops working or works differently Keyword optimization. In classical search, the marketer built a semantic core and worked to make a page maximally relevant to a particular phrase. In the answer environment, the user does not enter a keyword, but asks a question. And that question may be long, nuanced, and full of context and constraints. Google describes the technique used by AI Mode as “query fan-out”: the system breaks a single user question into subtopics and simultaneously looks for information on each of them [4][6]. That means a page perfectly tuned to one phrase may fail to appear in any of the fan-out subqueries if it does not cover adjacent aspects of the topic. The fight for position in a list of links. In the world of ten blue links, position one was the ultimate goal. In the answer environment, there is no position as such. There is the fact of being present inside the synthesized answer—and there is the role the brand plays inside it: the brand may simply be mentioned, may be cited, or may define the frame of comparison itself. The data shows that there is a correlation between classical ranking and citation in AI answers, but it is far from linear. According to AirOps, pages in Google’s first position are cited by ChatGPT in 43% of cases—3.5 times more often than pages outside the top 20 [7]. But that also means that 57% of first-position pages are not cited at all. The relationship exists, but it is not automatic. Traffic as the primary metric of success. If 69% of search queries end without a click, and among queries with AI Overview CTR falls to 0.6%, measuring success only through site visits means seeing an ever smaller portion of the picture. In the answer environment, a brand can shape a user’s decision without receiving a single visit to its site. That does not mean traffic has become unimportant. It means traffic has ceased to be the only currency. Seer Interactive found an important asymmetry: when a brand is cited inside an AI Overview, organic CTR is 35% higher than for uncited competitors on the same queries [3]. So citation inside the answer does not kill traffic—it redistributes it in favor of those the system regards as sufficiently trustworthy sources. Content for volume. Many SEO strategies in recent years were built on large-scale content production: the more pages covered the semantic core, the broader the reach. In the answer environment, that logic breaks down. An answer system does not sift through hundreds of site pages looking for the answer; it selects the best fragment from the best source. Ten weak articles on one topic lose to one strong article—not because the algorithm “penalizes” quantity, but because when synthesizing an answer, the system selects the most convincing and verifiable source. ## Where familiar optimization can backfire The academic GEO study directly showed that traditional SEO tactics such as keyword stuffing performed weakly in a generative context [5]. But the harm can be less obvious than that. The first risk is the language of self-presentation instead of the language of the task. Companies accustomed to SEO often describe themselves in the language of their own marketing categories: “leading platform,” “comprehensive ecosystem,” “innovative solution.” An answer system operates in the user’s language, and the user asks differently: “what should I choose for a small store,” “how is one solution different from another,” “what is better if the budget is limited.” If a brand has spent years optimizing content around its own terminology rather than the language of real demand, AI may simply fail to connect it to the user’s task. The second risk is an excess of self-description without external confirmation. In the SEO world, a strong site could dominate branded queries while relying mainly on its own content. In the answer environment, the system looks for external confirmation. If a brand claims certain advantages but no independent source confirms them in other words, the answer system will be more cautious in recommending it. A strong SEO site without an external trust contour is a vulnerable structure. The third risk is technical barriers for AI crawlers. Some sites that are used to finely managing crawl budget block certain crawlers through robots.txt. According to Press Gazette, around 80% of major news publishers already block at least one crawler from an AI system [8]. For media companies protecting their content, that is a deliberate choice. But for a commercial brand that wants to be visible in answers, blocking can amount to voluntarily leaving the field. ## The new task: not ranking position, but the right to be cited Taken together, this produces a picture in which SEO is not dying, but its task is being transformed. The goal used to be to reach the first page of results. The goal now is to become a source an AI system will want to rely on when forming an answer. That is a more difficult position, but also a more valuable one. To do that, a company has to do several things that classical SEO did not always require. First, write in the language of the task rather than the language of the brand. If the user asks, “which service is suitable for a small team without an analyst,” and the brand’s site says “modular environment for intelligent data management,” the connection will not form. Second, support claims with concrete data and external links. The Princeton study showed that adding statistics increases the probability of citation by an AI system by 30–40% [5]. Third, build not only the site, but the entire source contour: external reviews, case studies, comparison materials, industry mentions, and presence in knowledge graphs. Fourth, rethink the metrics: alongside traffic, there should be citation share, frequency of appearing on the shortlist, and the quality of the brand’s role in the answer. For a small business owner, this may sound intimidating. But in practice, many of these actions do not require a huge budget. They require clarity: who you are, what you do better than others, who can confirm it, and what language your customer uses to describe the task. This is not a question of technical optimization, but of intellectual honesty in describing your own brand. --- ## What seems well established It seems well established that the technical foundation of SEO (indexability, speed, structured data) remains a necessary condition for appearing in AI answers. It is also well established that traditional tactics such as keyword stuffing do not work in a generative context, while specificity, verifiability, and external authority meaningfully increase the probability of citation. ## What remains uncertain or platform-dependent What is less firmly established is the precise degree of correlation between classical ranking and citation in answer systems beyond Google AI Overviews. For ChatGPT, Perplexity, and Copilot, that relationship has been studied less and, according to preliminary evidence, appears to work differently. ## Practical implications for brand work For a company, this means the SEO team needs to broaden its field of view: continue building the technical foundation, but stop treating position in a list of links as the end goal. The new goal is to become a source AI relies on when forming an answer. And that requires not only a site, but the full contour of confirmation around it. --- ## Sources [[1] Similarweb. Zero-Click Search Research: 56% to 69% Growth (May 2024 – May 2025). 2025](https://www.similarweb.com/blog/marketing/geo/what-is-geo/) [[2] Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results. 2025](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) [[3] Seer Interactive. AIO Impact on Google CTR: September 2025 Update. 2025](https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-september-2025-update) [[4] Google Search Central. AI Features and Your Website. 2026](https://developers.google.com/search/docs/appearance/ai-features) [[5] Aggarwal P., Murahari V., Rajpurohit T., Kalyan A., Narasimhan K., Deshpande A. GEO: Generative Engine Optimization. KDD '24, ACM, 2024](https://dl.acm.org/doi/10.1145/3637528.3671900) [[6] Google Search Help. Get AI-Powered Responses with AI Mode in Google Search. 2026](https://support.google.com/websearch/answer/16011537) [[7] AirOps. Citation Analysis: SERP Position vs. ChatGPT Citations. 2026](https://www.position.digital/blog/ai-seo-statistics/) [[8] Press Gazette. Nearly 80% of top news publishers now block at least one AI training crawler. 2025](https://www.frase.io/blog/what-is-generative-engine-optimization-geo) --- Source: https://ansmeter.com/knowledge-base/practical-action-map # Practical action map: how to strengthen a brand’s machine distinctness **Research question.** What specific actions can a company take to improve its brand’s machine distinctness, and in what sequence do those actions have the greatest effect. **Evidence type.** Google Search Central documents, the academic GEO study (Princeton, Georgia Tech, IIT Delhi), empirical data from Search Engine Land, AirOps, and Seer Interactive, and market observations from Similarweb and McKinsey. **Data freshness.** The recommendations are based on data and practices current as of March 2026. ## Why this requires a map, not a list of tricks When a company first learns that its brand is poorly visible to AI, the first reaction is almost always the same: “What exactly do we need to redo on the site?” It is an understandable reaction, but it leads into a trap. The problem of AI visibility almost never comes down to a single page, a single paragraph, or a single technical fix. It is distributed across several layers, and that is precisely why what is needed is not a scatter of tips, but a map: a sequence of actions in which each step strengthens the effect of the previous one. This article is structured as that kind of map. It moves from the simplest to the more complex, from internal changes to external ones, from quick actions to long-term ones. Each step is tied to existing articles in the Ansmeter corpus, so you can go deeper wherever it becomes useful. The map does not promise miracles, but it does help you understand where to begin and how not to waste effort. ## Step 1. Check identity: can the machine identify you correctly Before thinking about visibility, it is worth making sure the answer system can distinguish your brand from others at all. That sounds obvious, but in practice this is exactly where problems begin. If a company uses several names, if the legal name diverges from the product name, if different pages describe the brand in different words, the machine gets not a stable entity but a set of partially overlapping signals. What to check right now: does the company have the same name on the homepage, in the documentation, in press releases, in profiles on external platforms, and in structured markup? Does the category description on the site match the way users describe it? Does the brand have an entry in Wikidata, Google Knowledge Graph, or at least the main industry directories? A company’s presence in knowledge graphs and directories increases the probability of correct entity identification in the answer. This is not magic — it is simple logic: if the machine can verify that “Company A” is exactly that company, and not another one with a similar name, it will recommend it with greater confidence. *Related material: [Why a strong brand can be invisible to AI systems](/knowledge-base/why-strong-brands-become-invisible) — a detailed breakdown of the five layers of machine distinctness.* ## Step 2. Rebuild the language: speak not about yourself, but about the user’s task This is perhaps the most important and the most underrated step. Most websites are written in the language of the brand: “we are a leading platform,” “our ecosystem of solutions,” “an innovative approach to data management.” An answer system operates in a different language — the language of the task the user is trying to solve. A person does not ask, “show me an ecosystem of solutions,” but rather, “what should I choose for a store with a small team if I do not want a long implementation.” The GEO study from Princeton, Georgia Tech, and IIT Delhi showed that among the nine tested optimization strategies, the best results came from those that increase specificity and verifiability: adding statistics, citing authoritative sources, and including expert quotations [1]. All three are ways of moving from abstract self-description to concrete language the machine can extract and use. What to do: collect 20–30 real questions customers ask during the buying process (from sales, reviews, and forum discussions). Compare them with the language the brand uses to describe itself on the site. If the overlap is weak, rewrite the key pages so they answer those questions directly. *Related material: [Category drift](/knowledge-base/category-drift) — shows how a brand loses when AI translates the task into someone else’s language.* ## Step 3. Strengthen structure: make the content extractable for the machine An answer system does not read a site the way a person does — from the first paragraph to the last. It extracts fragments: an answer to a question, a definition, a comparison, a fact with a number. If your page is a long marketing text with no clear headings, no definitions, and no facts, the machine has nothing to extract from it. According to Search Engine Land, more than 82% of the pages cited in Google AI Overviews are “deep” content pages (two or more clicks away from the homepage), not the homepages themselves [2]. That makes sense: deep pages usually contain specifics — product descriptions, comparisons, instructions, and case studies. That is exactly the type of content an answer system can most easily extract and reuse. What to do: on each key page, make sure the first 40–60 words give a direct answer to the question that page is meant to cover. Use question-based headings (H2, H3). Add concrete numbers, examples, and comparisons. Implement structured markup — at a minimum Organization, Article, FAQ, and Product where applicable. *Related material: [Machine-readable commerce infrastructure](/knowledge-base/machine-readable-commerce-stack) — a detailed analysis of the data and markup layer.* ## Step 4. Build an external trust contour: give the machine a way to verify your claims This is the step that separates strong AI visibility from average AI visibility. Your own site explains what the brand would like to be considered. External sources show what it is actually considered to be. The answer system tries to reconcile those two versions — and if there is no external validation, it will be more cautious in recommending the brand. According to Search Engine Land, around 85% of brand mentions in AI answers come from third-party pages rather than from company-owned websites [2]. That does not mean your own site is unimportant. But it does mean that without an external trust contour, even an excellent website is not enough. What to do: build a map of external sources that already mention the brand (reviews, directories, industry media, Reddit, YouTube). Identify which key brand claims are independently validated and which exist only on the brand’s own site. Work on the gaps deliberately: offer expert commentary, publish guest articles, and secure presence in category lists and comparisons. Do not forget Wikidata and industry directories. *Related material: [External authority versus the brand’s own website](/knowledge-base/external-authority-vs-own-site) — examines which sources actually shape a brand’s right to be recommended.* ## Step 5. Ensure technical accessibility for AI crawlers Answer systems rely on web search and external document retrieval. If a crawler cannot crawl the site, the content will never make it into the answer context. Google Search Central emphasizes that for a page to become a source for AI Overviews or AI Mode, it has to be indexed and allowed to appear with a text snippet [3]. What to check: are answer-system crawlers blocked in robots.txt (OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended)? Are the XML sitemap and IndexNow working properly (for Bing/Copilot)? Does the content load without JavaScript or via server-side rendering — since many AI crawlers do not execute client-side JS? *Related material: [Access economics](/knowledge-base/access-economics) — describes four modes of content access and helps determine an access policy.* ## Step 6. Start observing: measure not only traffic, but citations too The final step turns a one-time fix into a system. Without observation, you cannot understand what is working, what broke, or where the next actions are needed. The minimum observation set even a small team can maintain is this: once a month, ask 10–20 key questions in your category in ChatGPT, Google AI Mode, and Perplexity. Record whether the brand appears, in what role (mentioned, cited, defining the frame), and which competitors are named above it. For structured tracking, you can use the mini research card — a template that already exists in the Ansmeter corpus. A more mature team can add automated tracking through tools such as Similarweb AI Search Intelligence, Semrush AI Visibility Toolkit, or Bing Webmaster Tools AI Performance. But even a manual run of 20 questions once a month already gives more data than no observation at all. *Related materials: [Mini research card](/knowledge-base/mini-research-card) — a template for recording observations. [Mention, citation, and influence](/knowledge-base/mention-citation-influence) — the three levels of presence worth distinguishing in observation work.* ## Sequence of actions and realistic timing The ideal order is exactly the one described above: from identity to language, from language to structure, from structure to the external trust contour, from that contour to technical accessibility, and from accessibility to observation. Each layer creates the foundation for the next. For a small company with a single marketer, the realistic horizon for the first three steps is 4–8 weeks. The external trust contour is a slower process and deserves a 2–3 month window. Observation begins immediately and never ends. It is important to remember that between a change in a brand fact and its stable appearance in an AI answer, time passes — the update lag described in a separate corpus article. So there is no point expecting an immediate result. The right sequence of actions, plus patience, plus regular observation creates a cumulative effect that strengthens over time. --- ## What seems well established It seems well established that specificity, verifiability, and external authority significantly increase the probability that a brand will be cited in answer systems. It is also well established that technical site accessibility for crawlers remains a necessary condition, and that a company-owned website without an external contour of validation is insufficient for stable visibility. ## What remains uncertain or platform-dependent What is less firmly established is the exact weight of each factor across different platforms and industries. The optimal combination of actions depends on the category, the language, the query type, and the brand’s current level of presence. ## Practical implications for brand work For a company, this means that work on AI visibility should be managed not as a one-off project, but as an operating discipline: from identity audit through rebuilding language and the trust contour to regular observation of the outcome. --- ## Sources [[1] Aggarwal P., Murahari V., Rajpurohit T., Kalyan A., Narasimhan K., Deshpande A. GEO: Generative Engine Optimization. KDD '24, ACM, 2024](https://dl.acm.org/doi/10.1145/3637528.3671900) [[2] Search Engine Land. AI Overview citation analysis: 82.5% deep content, 85% third-party sources. 2025](https://www.gen-optima.com/blog/how-to-improve-brand-visibility-in-ai-search/) [[3] Google Search Central. AI Features and Your Website. 2026](https://developers.google.com/search/docs/appearance/ai-features) [[4] McKinsey & Company. New front door to the internet: Winning in the age of AI search. 2025](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search) [[5] Seer Interactive. AIO Impact on Google CTR: September 2025 Update. 2025](https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-september-2025-update) --- Source: https://ansmeter.com/knowledge-base/field-note-category-language-gap # Field note from a research run: how site language made a brand invisible in its own category > Field note — a short format that captures one concrete observation from a real Ansmeter research run. It does not claim to be a complete study; it shows how the theoretical concepts in the corpus appear in practice. --- ## What we saw In a research run on the category “analytics for mid-market e-commerce,” one of the brands being tested consistently landed in 4th–5th position in ChatGPT answers and barely appeared in Perplexity. This was unexpected: the brand is not small, it has a strong site with detailed documentation, an active blog, and several case studies with major clients. In traditional Google search, it ranks in the top 5 for its core category queries. By all the usual measures, this is a visible brand. But in neutral Ansmeter scenarios, where the models were asked questions such as “what should I choose for a mid-sized online store with a small team,” the brand lost to competitors that were objectively less well known. ## What turned out to be the cause The analysis showed three overlapping problems. The first and most important was a gap between the language of the site and the language of demand. The brand described itself as a “modular environment for intelligent commerce analytics.” Users — and models reflecting their language — were asking about a “simple reporting service for an online store without a dedicated analyst.” Those two languages barely intersected anywhere. Nowhere on the site was there a direct answer to the question “who is this for?” in the same words a user would use to describe the task. The second problem was that the model reformulated the task into an adjacent category. Instead of “analytics for e-commerce,” the answer was built around “BI tools for small business.” As a result, the list included two solutions from a neighboring category that the brand did not consider competitors. This is the classic category drift described in a separate article in the corpus. The third problem was the update lag. Two months before the research run, the brand had launched a new pricing plan aimed at the mid-market segment. But the model still described it as a solution for large companies — information about the new plan had not yet seeped into external reviews or structured data. ## What follows from this The observation confirms several theses that the Ansmeter corpus describes as persistent. Machine distinctness and human recognizability are different things. A brand that is well known to people can be functionally invisible to a model if its language does not match the language of the task. Category drift happens before direct comparison between brands. The brand lost not to a competitor, but to someone else’s frame. The model first renamed the task and only then assembled a list within the new category. Update lag is not an abstract delay. It is a concrete situation in which the brand has already changed a fact about itself, but the machine has not yet had time to see it. ## What the brand could have done Add a page to the site that directly answers the question “who is our product for?” in the language of the user’s task, not the internal marketing category. Make sure the new pricing plan is described not only on the pricing page, but also in external reviews. Check that the structured markup (Product, Offer) reflects the current pricing plans and target audience. A repeat research run in 6–8 weeks would show whether the picture changed. --- Source: https://ansmeter.com/knowledge-base/language-geography-visibility # Visibility through the lens of language and geography **Research question.** Why can the same brand look substantially different in AI answers across different languages and countries, and what practical implications follow from that. **Evidence type.** Google data on the expansion of AI Mode to ~100 languages, Google Search Central documentation, Similarweb and Semrush studies on cross-platform visibility, and SparkToro data on the instability of recommendations. **Data freshness.** The platform data is current as of the first quarter of 2026. > There is no single AI visibility even within one platform. Change the language of the query or the user’s geography, and the brand can move from the top three to the periphery of the answer — or disappear entirely. ## The illusion of unified visibility The Ansmeter corpus already contains an article about the “answer bubble” — the effect in which a brand looks different in ChatGPT, Google, Copilot, and Perplexity. But there is another layer of divergence that companies notice less often: the same brand, the same platform, but a different query language — and the answer changes radically. A company that confidently lands in ChatGPT’s top three recommendations in English may not be mentioned at all in the answer to the same question in Russian, Arabic, or Japanese. This is not a bug. It is a consequence of how modern AI systems are built. A language model is trained on corpora of text, and the distribution of those texts across languages is extremely uneven. English dominates the training data of most major models. That means that for English-language queries, the model has a denser and more up-to-date map of entities, properties, and relationships. For less represented languages, that map is sparse: some brands are present in it, others are not, and still others are present but with distorted attributes. ## The scale of language expansion The problem stopped being marginal once answer systems moved beyond English at mass scale. By May 2025, Google AI Overviews were operating in more than 200 countries and 40 languages [1]. AI Mode, launched in March 2025 in English only, expanded to nearly 100 languages by February 2026 through three major waves: November 2025 (35 languages), February 2026 (53 languages) [2]. ChatGPT serves more than 900 million weekly active users around the world, a significant share of whom do not work in English [3]. For an international brand, this means AI visibility has to be checked not in one language, but in every language in which its markets operate. One run in English says nothing about what the brand looks like in Spanish, German, or Chinese. ## Three mechanisms of language divergence The first mechanism is asymmetry in training data. A brand that is heavily described in English-language media, reviews, and directories may be only weakly represented in the Russian-language or Arabic-language segment of the internet. The model knows it in one language, but literally does not “remember” it in another. This is not a translation issue — it is an issue of presence in the sources of a specific language segment. The second is different web retrieval behavior. When an answer system uses web search to supplement an answer — as ChatGPT Search, AI Overviews, and Perplexity do — it looks for sources in the language of the query. If the brand has no high-quality content in that language, the system will find competitors that do. The user will receive a recommendation without your brand — not because you are worse, but because you are absent in that language. The third is category drift through language. The same product can belong to different categories in different language cultures. What is described in English as an “analytics platform” may be interpreted in Russian as a “reporting service” or a “business intelligence system” — and each of those terms pulls in a different set of competitors and comparison criteria. The model inherits these categorical differences from the training data. ## The geographic layer Language is not the only variable. Geography changes the answer too. Google AI Mode and AI Overviews take the user’s location into account when forming an answer [1]. Perplexity retrieves web sources with regional relevance in mind. That means the same English-language query from London and from Singapore may produce different results — with different competitors, different prices, and different recommendations. For brands with several regional versions of a site, this creates an additional layer of complexity. You need to check not only “are we visible in English,” but also “are we visible in English from the United Kingdom,” “are we visible in English from India,” and “are we visible in German from Germany and from Switzerland.” ## Instability as a baseline property The problem is made worse by the fact that AI-system answers are stochastic by nature. According to SparkToro, the probability that ChatGPT or Google AI, across 100 repeated queries, will produce an identical list of brands in even two answers is less than 1% [4]. Every answer is a probabilistic sample, not a deterministic result. Changing the language and geography adds another layer to that stochasticity. For diagnostics, this means that one run is not enough even within a single language. And for an international brand, the minimum meaningful diagnostic requires a matrix: several languages × several platforms × several repeats. ## What follows from this in practice For a company operating across multiple markets, language and geographic diagnostics are not a luxury, but hygiene. At a minimum, the following makes sense: Check visibility in every language in which the brand wants to be visible. Do not assume that a strong position in English carries over into other languages. Create content in the target language rather than relying on automatic translation. Models are trained on human texts, and the language of machine translation often does not match the language of real demand in a given culture. Make sure the external trust contour — reviews, directories, and industry media — covers the target languages. A brand with 50 English-language reviews and zero Russian-language reviews will be invisible to the model when the query is in Russian. Add multilingual labels in Wikidata. This is one of the fastest ways to give AI systems an entity anchor across several languages. --- **What seems well established:** Answer systems produce substantially different answers across languages because of asymmetry in training data, differences in web retrieval, and categorical divergence. Google AI Mode already supports ~100 languages, and AI Overviews support 40+, which makes language diagnostics relevant for any international brand. **What remains uncertain or platform-dependent:** The exact degree of divergence between language versions of answers is still poorly studied and depends on the category, the brand, and the platform. There are still few systematic comparative studies across languages. **Practical implications for brand work:** An international brand needs to check AI visibility in every target language and every target geography — and not rely on the results of a single English-language run. **Sources:** [1] Google Search Central. AI Features and Your Website. 2026 [2] ALM Corp. Google AI Mode Expands to 53 Languages: Complete Analysis. 2026 [3] OpenAI. Scaling AI for Everyone. 2026 [4] SparkToro. AI recommendations inconsistency: <1% chance of repeat brand list. 2026 --- Source: https://ansmeter.com/knowledge-base/multimodal-visibility # Multimodal distinctness: when a brand is searched not with words **Research question.** How do visual search, voice queries, and multimodal interfaces change the requirements for brand visibility, and what from classical text optimization carries over into the world of images, voice, and video. **Evidence type.** Google data on Google Lens (20 billion queries per month), Google documentation on multimodal AI Mode, and market observations from Semrush and Lumar. **Data freshness.** Platform and query-volume data are current as of the first quarter of 2026. > Search is no longer purely text-based. Google Lens processes 20 billion visual queries per month, AI Mode accepts photos as input, and voice queries are becoming longer and more contextual. A brand optimized only for text is losing a growing share of the audience. ## Text is no longer the only input Throughout the Ansmeter corpus, we have discussed visibility in the context of text queries: the user types a question, and the model produces an answer. But the world of search has long ceased to be reducible to typing words on a keyboard. A user photographs a product in a store and asks, “How much is this online?” They say aloud, “What model is this?” while pointing the camera at a pair of headphones. They upload a screenshot from Instagram and ask, “Find something similar, but cheaper.” They record a video and add a text question: “What material is this made of?” These are not exotic scenarios. Google Lens processes more than 20 billion visual queries per month, and 20% of them are shopping-related [1]. AI Mode is integrated with Google Lens: a user can take a photo or upload an image, and the system, using Gemini’s multimodal capabilities, analyzes the entire scene—objects, their context, materials, colors, and shapes—and produces a composite answer [2]. ChatGPT with GPT-4o processes images, voice, and text simultaneously. 27% of mobile users already use voice search [3]. For a brand, this means that text optimization is a necessary but no longer sufficient condition for visibility. If your product cannot be recognized from a photo, if your YouTube video has no transcript, if a voice assistant cannot link the spoken company name to the correct entity, you lose the audience that searches without words. ## How visual search changes the rules Visual search works in a fundamentally different way from text search. The user does not describe what they are looking for—they show it. Convolutional neural networks (CNNs) transform an image into a numerical vector and compare it with a database of indexed images [4]. This means that the quality, consistency, and technical accessibility of images on a site directly affect whether your product will be found. For e-commerce, the consequences are most obvious. A shopper sees a dress on the street, photographs it, and Google Lens shows similar products with prices from different online stores within three seconds. If your product images are low quality, lack descriptive alt text, lack Product schema, and do not follow a consistent photography style, they will not make it into that selection. A competitor with clean, marked-up photos will. Visual consistency across platforms is also becoming a factor. Google Lens is better at recognizing brands that use the same photography style across their website, marketplaces, and social platforms. A fragmented visual layer makes it harder to tie content to an entity [5]. ## Voice search and long queries Voice queries differ from text queries not only in modality, but also in structure. A person speaking aloud uses natural sentences: “What’s the best café near me that’s open right now?” instead of “cafe near me open.” Queries in AI Mode are, on average, three times longer than ordinary search queries [6]. This means that content optimized for short keyword phrases may not match the way people formulate queries by voice. For a brand, the practical implication is clear: FAQ sections written in a question-and-direct-answer format work better for voice search than long marketing texts. Structured data (FAQ schema, HowTo schema) helps voice assistants extract a specific answer. The brand name should be pronounceable and unambiguous—a model that cannot connect the spoken “Exco-Data” to the entity “ExcoData” will lose the brand in a voice query. ## Video and transcripts AI systems are making increasing use of video content. YouTube video transcripts become a source for citation: if an expert in your video explains in detail how the product works, and the transcript is available, the model can extract a passage from it for an answer. If there is no transcript, the video remains invisible to the text layer of the answer system. Google explicitly states that AI Mode uses multimodal analysis: the system works simultaneously with text, images, video, and context [2]. For a brand that publishes educational videos, reviews, or product demos, a clean and accurate transcript is not optional—it is a condition of discoverability. ## What to do right now Multimodal optimization does not require a revolution. It requires extending familiar work into new formats. Images: high quality, descriptive filenames and alt text, Product schema tied to specific products, and a consistent photography style across platforms. Voice: FAQ sections in question-answer format, HowTo schema for instructions, and a pronounceable, unambiguous brand name. Video: transcripts for every video on YouTube and on the site, VideoObject schema, and descriptive titles and metadata. General layer: the same principle as for text visibility—structured data, machine readability, and external confirmation. Multimodality does not replace these foundations; it adds new input channels on top of them. --- **What seems well established:** Visual search already processes tens of billions of queries per month. AI Mode integrates multimodal input (photo + text + voice). Video transcripts are used as a source for citation. Voice queries are longer and more conversational than text queries. **What remains uncertain or platform-dependent:** The exact share of AI answers initiated by visual or voice input is still poorly measured outside Google Lens. The effect of multimodal optimization on brand citation across different platforms has been studied only fragmentarily. **What this changes in practice:** A brand needs to optimize not only text, but also images, video, and voice discoverability. The basic actions (alt text, transcripts, FAQ schema) are simple and can be started right away. **Sources:** [1] Google / DemandSage. Google Lens: 20 billion visual searches per month, 20% shopping-related. 2025 [2] 9to5Google / Google I/O. Google AI Mode adding multimodal Google Lens search. 2025 [3] Google / Lumar. 27% of global mobile users use voice search. 2025 [4] Xictron / Pinecone. Visual search technology: CNN embeddings and vector matching. 2026 [5] SE Blog. Multimodal Search Optimization: visual consistency and entity recognition. 2026 [6] ALM Corp. Google AI Mode queries average nearly 3x longer than traditional search. 2026 --- Source: https://ansmeter.com/knowledge-base/platform-change-chatgpt-instant-checkout # ChatGPT Instant Checkout: purchasing without leaving the conversation > Dated platform change — a short format that captures a specific platform change and its significance for a brand’s visibility in AI. It does not aim to be a comprehensive review; it records the fact and helps the reader quickly understand what happened. **Date of change:** September 2025 (launch), January 2026 (expansion through Shopify) **Platform:** ChatGPT (OpenAI) **What happened:** OpenAI launched Instant Checkout, a feature that allows users to make purchases directly inside a ChatGPT conversation without navigating to an external website. The first partners were Etsy and Shopify merchants (Glossier, SKIMS, Spanx, Vuori). Payments are processed through Stripe. OpenAI takes a 4% commission on each transaction. **Why it matters for visibility:** The entire choice chain — from question to purchase — now takes place inside ChatGPT. The brand’s website receives no visit. Analytics sees no traffic. The purchase decision is made on the basis of the data the model was able to extract and present inside the conversation. This means that brand visibility inside a ChatGPT conversation becomes not only a matter of reputation, but a direct sales factor. A brand that does not make it into the recommendation does not simply lose a mention — it loses the transaction. **What to check:** - Whether your product appears in catalogs accessible through Shopify or Etsy. - Whether your product’s structured data is accurate (name, price, availability, description, images). - How ChatGPT describes your brand when asked directly about the category — and whether you appear in the recommendation. --- Source: https://ansmeter.com/knowledge-base/agentic-choice # When the buyer is not a person but their agent **Research question.** How brand visibility changes when an autonomous AI agent stands between the company and the buyer, searching, comparing, and making decisions on its own. **Evidence type.** Salesforce data from Cyber Week 2025, OpenAI and Google documents on agentic protocols, Gartner and McKinsey forecasts, and market observations from NRF and MIT Sloan. **Data freshness.** The platform data and market forecasts reflect the state of play in the first quarter of 2026. ## From answer to action Throughout the Ansmeter corpus, we have been talking about a world in which the user asks a question and the system forms an answer. That is the world of the answer environment: the brand competes for the right to be named, cited, and recommended. But the next wave is structured differently. An AI agent does not merely answer a question — it acts: it searches for products, compares options, checks prices and availability, and in some cases already completes the purchase without showing the person an intermediate list of alternatives. This is not a distant future. ChatGPT Instant Checkout has been operating since September 2025, allowing users to make purchases directly inside the conversation through Shopify and Etsy partners [1]. In January 2026, Google announced its own Universal Commerce Protocol (UCP), joined by Walmart, Target, and more than 20 partners [2]. Amazon is building a closed ecosystem — Rufus AI and Alexa+ — inside its own platform. Three competing systems are already live, and none of them asks brands for permission to participate. ## A scale that can no longer be ignored Salesforce data from Cyber Week 2025 shows that the volume of tasks completed by AI agents on shoppers’ behalf grew 70% versus 2024, and 84% on Black Friday [3]. Gartner forecasts that by the end of 2026, 25% of enterprise software procurement will involve an AI agent [4]. McKinsey estimates that personalization and autonomous AI-assisted purchasing could create an additional $1.2 trillion in value for global retail [5]. According to Human Security, agentic AI activity grew 6,900% over the past year, and 64% of consumers planned to use AI for holiday shopping in 2025 [6]. For a brand, these figures point to a simple but painful reality: one of the most important “buyers” in the near future may not be human at all. And if it is not human, then website design, emotional positioning, visual identity, and banner ads are irrelevant to it. What it needs is structured data it can read, compare, and verify. ## Three ecosystems, three different languages Today, agentic commerce is taking shape around three incompatible protocols. ACP (Agentic Commerce Protocol) is the OpenAI–Stripe protocol that works through ChatGPT Instant Checkout. OpenAI takes 4% from each transaction plus the standard payment processor fee [1]. UCP (Universal Commerce Protocol) is Google’s coalition protocol, connected to AI Mode and Gemini, with support from Walmart, Target, and Shopify [2]. Amazon’s closed ecosystem — Rufus AI, Alexa+, and Buy for Me — works only inside the Amazon marketplace. For a brand, that means it must be visible in all three worlds. A company that has optimized its data for only one platform loses demand on the other two. And none of these worlds operates by the rules of classic SEO. ## What the agent sees — and what it does not An AI agent does not read marketing copy the way a human does. According to MIT Sloan, 82% of executives cite data quality as the main barrier to achieving generative AI goals [7]. The agent works with structured attributes: name, category, price, availability, specifications, return terms. If that data is missing, inaccurate, or out of sync across platforms, the agent cannot compare the product correctly — and it chooses the competitor with cleaner data instead. That is fundamentally different from the answer environment. In a model-generated answer, a brand can compensate for weak data with a strong reputation or vivid expert content. In an agentic scenario, those compensations disappear. The agent makes its decision on the basis of attributes, not impressions. If your product feed does not contain machine-readable data on compatibility, warranty, size, or delivery times, the agent may remove you from the comparison before it even gets to assessing quality. ## Analytics breaks down There is another unpleasant feature of the agentic world: classic web analytics stop working. When a user asks ChatGPT about a product category, the agent compares, recommends, and completes the purchase — all inside the conversation. The brand’s site gets no visit, no click, no session [1]. GA4 does not see that traffic. A marketer who measures success by site visits does not notice that the sale happened through another channel. That does not mean analytics has become useless. It does mean, however, that new metrics must appear alongside the familiar ones: citation in agentic recommendations, inclusion in agentic comparisons, conversion through agentic checkout. The tools for this are only beginning to emerge, but the need itself is already obvious. ## What this changes for the brand today Agentic choice has not yet become the dominant form of purchasing. But the pace of growth is such that within 12–18 months it could become a mandatory channel for any company selling goods or services online. And preparing for it does not begin with a new tool. It begins with fundamental work: clean structured product data, synchronized across platforms, machine-readable, and current. For a business owner, this means yet another shift in mindset. First, you had to learn to be visible to the search engine. Then, to be understandable to the answer system. Now you have to be usable for the agent — which does not ask, does not read, and does not click, but simply chooses on the basis of data. --- **What seems well established:** Agentic commerce is already real: ChatGPT Instant Checkout, Google UCP, and Amazon Rufus are processing real transactions. The volume of agentic tasks in retail is rising quickly, and the quality of structured data is becoming a critical factor in whether a brand is included in agentic recommendations. **What remains uncertain or platform-dependent:** The precise share of purchases made through AI agents is still small and highly category-dependent. The protocols are competing with one another, and it is not yet clear which one will become the standard. Metrics for agentic visibility are not yet settled. **What this changes in practice:** Brands need to prepare not only content for people, but also machine-readable data for agents: product feeds, structured catalogs, API-compatible product descriptions, and current attributes synchronized across platforms. **Sources:** [1] OpenAI / Opascope. ChatGPT Instant Checkout and Agentic Commerce Protocol. 2025–2026 [2] Google. Universal Commerce Protocol, January 2026 [3] Salesforce. Cyber Week 2025 Agent Data. 2025 [4] Gartner. Predictions for AI agent mediation in enterprise software. 2025 [5] McKinsey. AI-driven personalization and autonomous shopping value. 2025 [6] Human Security. Agentic AI activity growth: 6,900% YoY. 2025 [7] MIT Sloan. Data quality as a barrier to GenAI goals: 82% of executives. 2025 --- Source: https://ansmeter.com/knowledge-base/wikipedia-wikidata-knowledge-graph # Wikipedia, Wikidata, and the Knowledge Graph: the invisible foundation of AI visibility **Research question.** Why a brand’s presence in Wikipedia, Wikidata, and the Knowledge Graph has become a practical lever for AI visibility, and how to work with it. **Evidence type.** Analysis of ChatGPT citations (680 million citations, Semrush), Wikipedia traffic data, Google documentation on the Knowledge Graph, and market observations from Status Labs and LinkSurge. ## Why the encyclopedia became more important to the machine than the website When a company thinks about its visibility on the internet, Wikipedia usually does not make the priority list. That is understandable: a Wikipedia article seems secondary compared with the company’s own website, blog, advertising, or SEO. But for answer systems, the hierarchy looks very different. An analysis of 680 million ChatGPT citations from August 2024 through June 2025 showed that, among the top 10 most cited sources, Wikipedia accounts for nearly half — 47.9% [1]. This is not an accident. All major language models — ChatGPT, Gemini, Claude, Llama — were trained on corpora in which Wikipedia was intentionally given extra weight. The Google C4 dataset, one of the core training sets, deliberately increased Wikipedia’s share relative to other web sources [2]. And in June 2025, ChatGPT became Wikipedia’s top traffic referrer — creating a symbiotic loop in which AI cites the encyclopedia and users click back through to it [3]. For a brand, that means something concrete: if the company has a high-quality Wikipedia page, the answer system gets a reliable, neutral, verified source for entity identification. If there is no page, the model is forced to assemble information from less structured and less authoritative sources — and the result will be less accurate. ## Wikidata: the brand’s machine-readable passport Wikipedia is a text encyclopedia for people. Wikidata is a structured database for machines. Every entry in Wikidata has a unique identifier (Q-ID), which is used to anchor an entity unambiguously. Google Knowledge Graph draws directly from Wikidata [4]. When an answer system encounters a brand name, it first checks whether there is an entry for it in the knowledge graph — and that is where Wikidata becomes a critical link. Unlike Wikipedia, Wikidata does not impose strict notability requirements. A company that cannot get a Wikipedia article because it lacks sufficient media coverage can still create a Wikidata entry: specify the type of organization, industry, founder, products, and official website. That is enough to give the machine a stable identifier and a set of basic attributes. Brands without a Wikidata entry face a structural disadvantage. The answer system first checks whether the entity exists in the knowledge graph, and only then decides whether the site’s content is worth citing. If that check fails, the model will be more cautious in recommendations — or bypass the brand entirely [5]. ## Knowledge Graph: the map AI uses to navigate Google Knowledge Graph is not a standalone product but an infrastructure layer on which Knowledge Panel, AI Overviews, and AI Mode are built. It contains billions of entities and trillions of relationships among them. When a user asks a question, AI does not simply search for relevant documents — it first identifies entities through the knowledge graph and then selects sources for the answer. For a brand, that means inclusion in the Knowledge Graph is not a bonus but a foundation. Without it, the answer system has to spend additional compute resources just to understand who you are. Researchers call this a “comprehension budget”: the cheaper it is for the machine to identify your entity, the higher the probability of citation [5]. ## What to do right now Check whether the brand is present in Wikidata (wikidata.org). If there is no entry, create one with the basic properties: P31 (entity type), P452 (industry), P856 (official website), P112 (founder). This takes 15–30 minutes and requires no technical skills. If the brand meets Wikipedia’s notability criteria, prepare or improve the article. If it does not, do not force it: Wikidata already provides a basic level of identification. Make sure the site’s Schema.org markup (Organization, sameAs) points to the Wikidata Q-ID and other official profiles. That creates a closed identification loop that is easiest for the knowledge graph to verify. Maintain consistency: the brand’s name, description, and category should be the same in Wikidata, on the website, in Google Business Profile, and across all external directories. --- **What seems well established.** Wikipedia is the most cited source in ChatGPT and the second most frequent across all LLMs. Wikidata feeds directly into the Google Knowledge Graph. Brands with a Wikidata entry have a structural advantage when answer systems identify an entity. **What remains uncertain or platform-dependent.** The exact weight of Wikipedia and Wikidata relative to other trust signals varies by platform and is not fully disclosed. Having a Wikipedia page does not guarantee citation — the quality and freshness of the article also matter. **Practical implications for brand work.** Creating or improving a Wikidata entry is one of the fastest and least expensive ways to strengthen a brand’s machine identification. It is a “15 minutes of work with potentially long-term effects” kind of action. **Sources:** [1] Semrush / Status Labs. Analysis of 680M ChatGPT citations: Wikipedia at 47.9% of top-10. 2025 [2] Status Labs. How AI Models Use Wikipedia as a Truth Anchor. 2026 [3] ALLMO. Wikipedia-ChatGPT symbiotic loop: ChatGPT became Wikipedia's top referrer, June 2025 [4] Google. Knowledge Graph documentation; Wikidata as primary source. 2026 [5] LinkSurge. Entity Authority and AI Search Visibility. 2026 --- Source: https://ansmeter.com/knowledge-base/field-note-visibility-language-field # Visibility Language Field: why the same brand lives in different competitive worlds > A model can know a brand equally well in every language — and still recommend it differently. The language of the prompt determines not the accuracy of knowledge, but the competitive landscape in which the brand appears. **Research question.** Does a brand’s competitive set in AI model answers depend on the language in which the question is asked — and if so, to what extent. **Evidence type.** Ansmeter’s own research runs: five studies of one brand (Notion) across five languages, using GPT-5.4 and the standard scenario corpus. **Data freshness.** The runs were conducted on April 2–3, 2026. ## Five markets instead of one The article “Language and Geography of Visibility” examines the mechanisms through which language affects visibility. Here we look from the other direction — from data to conclusion: what exactly happens to a specific brand and its competitors when the language is switched. We chose Notion — a globally recognizable productivity product. Deliberately: we needed a brand the model clearly knows, so that we could rule out the explanation of “there just isn’t enough data.” Five runs, five languages, one model, one scenario corpus. Notion’s own score fluctuated moderately: from 62.9 in French to 75.7 in German, a spread of 12.8 points. The confidence intervals of the runs partially overlap. If we had stopped there, the conclusion would have been calm: “small fluctuations, possibly model noise.” But then we looked at the competitors. --- ## The competitor matrix: the main evidence | Brand | RU | EN | ES | FR | DE | Spread | |-------|----|----|----|----|-----|-------:| | **Notion** (target) | 71.2 | 68.8 | 69.1 | 62.9 | 75.7 | 12.8 | | **Slack** | 0.0 | 51.0 | 53.8 | 54.3 | 54.9 | 54.9 | | **Monday.com** | 47.4 | 30.5 | 29.0 | 7.8 | 13.0 | 39.5 | | **Asana** | 70.1 | 52.6 | 51.1 | 39.7 | 59.1 | 30.4 | | **Microsoft Copilot** | 36.2 | 39.9 | 42.8 | 51.1 | 24.8 | 26.3 | | **ClickUp** | 67.3 | 59.6 | 63.1 | 54.5 | 62.7 | 12.8 | | **Coda** | 46.0 | 46.7 | 38.4 | 41.1 | 43.6 | 8.2 | | **Airtable** | 33.5 | 37.5 | 30.5 | 28.3 | 40.4 | 12.1 | | **Confluence** | 22.3 | 13.2 | 16.6 | 24.0 | 20.6 | 10.8 | Notion’s spread is 12.8. Slack’s is 54.9. Monday.com’s is 39.5. Asana’s is 30.4. That is five different competitive landscapes compressed into one table. --- ## Three patterns we observed ### Binary disappearance: Slack In Russian, Slack receives a score of 0.0 — the model does not mention it at all in the context of productivity and workspace tools. In the other four languages, the result is stable: 51–55 points, with a spread of just 3.8 points. That kind of stability across four languages, combined with a complete zero in the fifth, is a strong argument that this is a stable property of the Russian-language field rather than a random outlier. The explanation most likely lies in the training data: Slack is actively discussed in English-, French-, and German-language sources as a team collaboration tool. In Russian-language sources, it is discussed hardly at all. The model did not lose knowledge of Slack; in this context, it never acquired it in the first place. ### Gradient of disappearance: Monday.com Monday.com shows a smooth decline from 47.4 in Russian to 7.8 in French. This is a third pattern, distinct both from Notion’s stability and from Slack’s binary switch. The brand seems to melt as it moves across language fields — retaining its presence, but losing weight. ### Inversion: Microsoft Copilot Where Notion is strongest (German — 75.7), Copilot is weakest (24.8). In French, the picture reverses: Notion is at 62.9, Copilot at 51.1. The two brands seem to sit on opposite ends of a seesaw, and language determines which one ends up higher. Based on our observations, this may be related to Microsoft’s activity in French-speaking European markets — but the data is not sufficient to make that claim with confidence. --- ## Knowledge is stable; recommendation is not An independent analysis of our data revealed a pattern that may matter even more than the competitor matrix itself. When the brand is already named in the prompt (diagnostic mode), the model answers about it with the same stability across all languages: 73–79 points, with a coefficient of variation of 3.7%. The model *knows* Notion equally well in Russian, French, and German. The divergence begins when the user has not yet named the brand. Average position in the answer, inclusion in the top three, citation of the notion.so domain — all of this depends heavily on language. In Russian, notion.so is cited in 24.5% of answers; in German, in 21.4%; in French, in 0%. For a brand, this leads to an uncomfortable conclusion: the model’s knowledge of you is a necessary but insufficient condition. The question is whether you make the shortlist before the user has spoken your name. The answer to that question depends on language. --- ## Three mechanisms we hypothesize We see three channels through which language reshapes the competitive field. All three are hypotheses supported by the data from these runs, but not experimentally validated. The first is asymmetry in training corpora. The model was trained on texts in which different brands are discussed with different frequency across different languages. Russian-language texts about productivity barely mention Slack; English-language texts mention it constantly. The second is different web sources. In web mode, the model searches in the language of the query and finds different reviews, comparisons, and rankings — with different brand mixes. French-language search returns French-language sources in which Notion is known, but notion.so is not cited. The third is different associative category graphs. In each language, the model builds its own map of the “productivity” category. In Russian, that map is Notion, Asana, ClickUp, Monday.com. In French, it is Notion, Slack, Microsoft Copilot, ClickUp. The mix of players differs, and that determines who ends up in the recommendation. --- ## What this means in practice For a category-leading brand, Visibility Language Field (VLF) is more a strategic task than a crisis. Your own score fluctuates moderately, but the competitors you are fighting in one language may be different in another. A strategy built around one competitive set risks becoming irrelevant in a neighboring language market. For a brand that is not dominant, the situation is harsher. Monday.com loses 40 points in the shift from Russian to French. Slack disappears entirely in Russian. If your brand occupies second or third place, VLF is a direct business risk: visibility earned in one language does not automatically transfer to another. The practical recommendation is simple: run a separate study in every target-market language. Compare not only your own score, but also the composition of the competitive set. A visibility growth strategy must account for the specific competitors that exist in each language — because across languages, those may be different companies. --- ## Methodological notes The data comes from five runs of one brand (Notion) on GPT-5.4. All runs used the standard Ansmeter corpus of 200 scenarios. Two runs (RU and FR) were conducted on April 2, and three (EN, DE, ES) on April 3, 2026. Differences between days (the day effect) were not separated from the language effect. The confidence intervals of the final scores partially overlap. Cochran’s Q test estimates the probability that the entire spread is explained by model noise at 4–8% — right on the edge of statistical significance. But the structural patterns — Slack’s stability across four languages with a zero in the fifth, the Monday.com gradient, the Copilot inversion — are poorly explained by stochasticity. The main limitation is clear: one brand, one model, one category, one run per language. Proper validation requires repeated runs (at least 5 per language) and tests on other brands and models. We call this observation VLF and consider it sufficiently well grounded for publication, but not sufficiently validated for final conclusions. ## What seems well established Changing the language of prompts reshapes a brand’s competitive set: some competitors appear, others disappear, and others radically change position. At the same time, the model’s diagnostic knowledge of the brand remains stable — what changes is the recommendation layer itself. ## What remains uncertain or platform-dependent One brand, one model, one run per language is not enough to draw conclusions about the scale of the effect in other categories and on other models. The exact boundary between the language effect and stochastic model noise has not yet been established. ## Practical implications for brand work An international brand needs to test visibility separately in every target-market language. The result of an English-language run does not carry over to other languages — especially for brands that are not unequivocal category leaders. ## Sources - [1] Ansmeter. Series of Notion runs across 5 languages: RU, EN, ES, FR, DE. GPT-5.4 model, corpus of 200 scenarios. April 2026. --- Source: https://ansmeter.com/knowledge-base/competitive-set-in-ai-visibility # Article 26. When Measurement Depends on the Neighbors: the Competitive Set in AI Visibility Studies *Overview and methodology article.* **Research question.** How does the composition of the brands being compared affect one brand’s visibility metrics in AI answers — and why does the competitive set need to be fixed if you want to compare runs honestly across time. **Evidence type.** Classic marketing literature on the competitive set (Aaker, Keller, Kapferer), formulas for AI visibility metrics in the Ansmeter methodology, and observations from our own repeated runs of one brand in April 2026. **Data freshness.** The observations and formulas refer to the April 2026 version of the Ansmeter methodology. > Replace one of the neighbors in the table — and your brand’s mention share rises or falls by ten points. The model’s behavior did not change; the competitive set within which that behavior was measured did. ## The same brand, two runs, different neighbors in the table In April 2026, we ran an audit for a moderately well-known SaaS brand twice, half an hour apart. The same models, the same prompt corpus, the same query language. No new players entered the category in those thirty minutes, the brand did not change its site or positioning, and even the AI providers’ caches remained almost fully warm between runs. The brand received almost the same overall score — a gap of about two points, well within the ordinary noise of repeated measurement. Only when we opened the second table — mention share within answers — did the picture change. In one model, that figure nearly tripled in half an hour. In another, it roughly doubled. In a third, the shift was more moderate, but still beyond the bounds of reasonable noise. The explanation turned out to be disappointingly simple: between runs, automatic competitor discovery slightly rewrote the list of neighbors. One player dropped out of the set, another took its place. The brand did not become more visible in answers. Its share was simply being calculated against a different group of neighbors — and arithmetic did the rest. This is not an anecdote about a broken system. It is a standard effect that appears whenever measurement pretends to measure one object while in fact measuring it within an environment of others. And almost any AI visibility tool measures in exactly that way. ## Where the term comes from “Competitive set” is a much older term than AI visibility. In Aaker’s and Keller’s work on brand equity, it is already a working concept by the late 1990s: a fixed circle of brands against which the focal brand is measured. In Kapferer, the closely related notion is the frame of reference: the circle of comparison within which a brand acquires meaning. The idea is simple: no brand exists by itself. It always has neighbors — the brands against which the buyer compares it at the moment of choice. And any researcher measuring brand strength has to name those neighbors explicitly. Otherwise, the numbers they obtain no longer have a clear referent. In a classic brand tracker, the competitive set is selected manually — usually two to four direct competitors plus one or two “indicator” players from adjacent categories. The list is fixed in the research design, and the same names are used again six months or a year later. If the researcher decides to update the set, that becomes a separate methodological decision and is noted in the report. In AI visibility, things usually work differently. Most often, the competitors are identified by the model itself: we give it a dozen or two prompts of the form “who else is worth considering in this category,” collect the answers, aggregate them, and obtain a set. This is convenient: the researcher does not need deep knowledge of another market and does not need to guess whom to include. But convenience has a cost. Automatic discovery produces a slightly different result from one run to the next. And that turns the set from a stable part of the design into a floating variable. ## Where the set hides inside the metrics AI visibility metrics fall into two classes, depending on how sensitive they are to the composition of the neighboring brands. It is useful to distinguish them, because they are built differently in principle, even though they sit side by side in a report and look similar. The first class consists of metrics that describe what happens to the brand itself. Did it appear in the model’s answer? In what position did it appear? How often did it make the top three? Did it receive an explicit recommendation? These indicators are calculated from the model’s behavior toward one brand. If the model answers in roughly the same way, the number stays roughly the same regardless of who else is in the set. The appearance of a new company in neighboring rows of the table barely changes them. The second class consists of share metrics. The brand’s mention share out of all mentions of all brands. The share of scenarios in which the brand surfaced, out of the total number of scenarios in which the model named anyone at all. The citation share of the brand’s domain out of all citations of competitors’ domains. These metrics are relative by nature. They have a numerator — what belongs to the focal brand. And they have a denominator — what belongs to the full set. If the set changes and the newcomer is mentioned less often than the player it displaces, the total denominator shrinks. The numerator does not move. The share rises. It is the same arithmetic by which you instantly look richer if the banker’s son leaves your class at school. Your own wealth has not changed, but your position in the distribution has — and any statistic that asks “what is your rank by income?” will now produce a different number. Nothing unfair has happened. The measurement simply depended on the composition of the group, and the group changed. That is why, when we say “the brand’s mention share increased,” we need to keep two different propositions in mind. The first is that the model really started mentioning the brand more often — behavior changed. The second is that the neighbors changed, the denominator was recalculated, and the share drifted — behavior may have stayed the same. Without an explicitly fixed set, the two cases are easy to confuse. And if a decision after the audit — whether to launch a campaign, shift positioning, or spend budget on content strategy — is made under the first interpretation when the second is the one that actually applies, the cost of that mistake can be high. ## Why automatic competitor discovery is always slightly different An AI model does not answer the same way to the same prompt every time. Even when the researcher drives randomness down as far as possible through generation settings, the model still has to choose among several plausible continuations, and that choice can differ slightly from run to run. That is not a defect; it is a property of how modern generative models work. When we ask, “who else is worth considering in category X?”, the answer almost always begins the same way — the same two or three undisputed leaders the model would return to anyone asking. The differences begin at the edge of the list. When it comes to the eighth, ninth, or tenth name, the model has several roughly equiprobable candidates in mind, and the ranking among them shifts slightly each time. If you aggregate a dozen or two such answers from one run and another dozen or two from a second, the aggregated lists will match at the top and diverge at the bottom. A brand that, the first time, received exactly enough votes to make the final eight will be ninth the second time and drop out. Another brand — one that was tenth last time — will take its place. From the standpoint of research design, this is bad news: the periphery of the set is mobile by construction, and no amount of system-side effort can fully stabilize it. You can increase the number of prompts through which competitors are identified — that helps, but not radically. You can warm provider caches — that reduces cost, but barely changes content. The jitter at the edge of the list remains. From the standpoint of practice, this means something simple: as long as the competitive set is redefined on every run, comparing runs to one another with any metric that includes a denominator is technically incorrect. The brand’s overall score will hold up, because it is built mainly from first-class metrics. Shares will not. ## What we do in Ansmeter Ansmeter solves this problem as follows. At the first audit of a brand, the system builds the competitive set in the usual way: automatic discovery plus the client’s option to add names they consider important or remove those they consider irrelevant. The final list — what the client approves before launch — is saved as the first version of that brand’s competitive set. At a repeat audit of the same brand, the system uses the same set by default. In the launch form, the client sees a prompt: “this brand already has a competitive set from [date], use it?” If the answer is yes, the repeat measurement is conducted against the same circle of neighbors as the previous one, and the share metrics become honestly comparable. If the client wants to revise the set — add a new player, remove one that has dropped out of the market, or refresh the list entirely — that becomes an explicit action that creates a new version of the competitive set. In the report, the line at the bottom of the methodology section shows which version was used, when it was created, and how many brands it contains. This is needed so that, when comparing reports, the client can see whether the comparison is between two runs with the same neighbor set or two runs with different ones. A fresh audit for a brand that has never been studied before discovers competitors from scratch — but the client can still intervene and adjust the set before launch. Here the system “remembers” nothing, because there is nothing yet to remember; but the formation of the first set remains under the client’s control rather than staying a hidden step. ## When it makes sense to update the competitive set In reality, there are not many situations in which an update makes sense. Six to twelve months have passed, and the category has changed visibly — someone has exited, someone has grown significantly, the market itself has shifted. In that case, the old set begins to misrepresent present reality, and it is worth refreshing it even if that breaks comparability with last year’s figures. Here, the value of an honest picture is higher than the value of continuous trend lines. A significant player has appeared who simply did not exist in the first set. If this means one or two names, it is easier to add them manually without launching a full revision — the overall structure of the set will remain, and most share metrics will stay comparable. But if many new names appear at once, that is probably a signal that the set should be refreshed in full. The brand has changed its positioning or shifted into an adjacent category. The competitive set should follow the brand — otherwise the measurement starts showing not real visibility, but visibility inside a category to which the brand no longer belongs. In all other cases, it is better to keep the old set. The natural temptation to “refresh it so it stays current” works against the usefulness of the research: every update resets the ability to compare with previous runs. Choosing to keep the set by default is methodological discipline in the purest sense, not conservatism. ## What seems well established AI visibility metrics whose formula contains a denominator over the entire brand set (mention share, scenario share, citation share) are recalculated when the composition of the set changes — even if the model’s behavior toward the focal brand does not change. The effect is built into the arithmetic of the metric, not into the system’s code; it will appear in any AI visibility tool that does not fix the set explicitly. ## What remains uncertain or platform-dependent The exact boundary between “a shift in shares caused by movement in the set” and “a shift in shares caused by a real visibility change” cannot be disentangled in a single repeat run. To separate them, you need either a fixed set, or two runs with the same shift in the set, or a dedicated stability analysis — a separate methodological task that applied AI visibility researchers still solve in different ways. ## Practical implications for brand work A repeat audit without a fixed competitive set shows not the brand’s dynamics, but the brand’s dynamics overlaid with shifts in the composition of its neighbors. For decisions based on comparing runs to one another — fix the set. For a fresh assessment of the market — update it. Do not confuse those two modes, and state explicitly in the report which mode the work is using. ## Sources - [1] Aaker, D. *Building Strong Brands.* The Free Press, 1996. - [2] Keller, K. L. *Strategic Brand Management.* 4th edition. Pearson, 2012. - [3] Kapferer, J.-N. *The New Strategic Brand Management.* 5th edition. Kogan Page, 2012. - [4] Ansmeter. Comparison of two consecutive runs of one brand (internal observation, April 2026). --- # Whom to read in AI ## Akari Asai Source: https://ansmeter.com/whom-to-read/akari-asai **Works on:** How a language model decides that its own internal knowledge is insufficient and that it should reach for an external source instead. Most retrieval-augmented systems are wired rigidly — always query the external store, then answer. Asai's Self-RAG (2023) showed that the model can be trained to make the decision itself: whether retrieval is needed here, what kind, and whether to trust what comes back. For Ansmeter this is a directly upstream line of work — modern AI engines in web-search mode already operate this way, and the decision to mention or omit a brand often happens at the retrieval layer rather than in the final generation. **Worth following when:** you need to understand why a model in web-search mode includes one source and ignores another on the same query. **Topics:** active retrieval in LLMs; adaptive RAG; the robustness of retrieval decisions to adversarial or noisy contexts. **Key works:** Self-RAG (2023, lead author); OpenScholar (2024, lead author); CRAG critic-corrective RAG (2024). --- ## Amir Globerson Source: https://ansmeter.com/whom-to-read/amir-globerson **Works on:** Using one language model as an adversarial interrogator of another to surface factual errors that neither could find alone. The 2023 "LM vs LM" paper from Globerson's group set up a cross-examination protocol: one model produces a claim, a second model asks follow-up questions designed to probe for inconsistencies, and the original claim is flagged if the follow-ups force a contradiction. The setup gives the field something it had been missing — a factuality-detection method that doesn't require labeled ground truth, doesn't require model internals, and doesn't require trusting either model on its own. Globerson's longer arc is in machine-learning theory, which shapes how the work reads: LLM evaluation as a problem about adversarial sampling and information geometry, rather than as a problem about prompt engineering. **Worth following when:** you want factuality detection methodology grounded in theoretical first principles rather than empirical recipe-hunting. **Topics:** adversarial LM-vs-LM evaluation protocols; factuality detection without labeled data; learning theory as a lens on LLM behavior. **Key works:** "LM vs LM: Detecting Factual Errors via Cross-Examination" (2023); broader ML-theory work informing the evaluation methodology. --- ## Ari Holtzman Source: https://ansmeter.com/whom-to-read/ari-holtzman **Works on:** How language models actually generate text from their probability distributions — and what the choice of sampling method does to evaluation results. The 2019 paper "The Curious Case of Neural Text Degeneration", which Holtzman led, introduced nucleus sampling (top-p): instead of sampling the next token from the full distribution or from the top-k highest-probability tokens, sample from the smallest set of tokens whose cumulative probability exceeds p. The method became the standard for almost every deployed LLM, including the ones Ansmeter currently evaluates, because it produced more fluent and less repetitive text than the alternatives. The implication for evaluation that the field has been slower to absorb: two runs of the same model with different sampling settings can produce noticeably different scores, which means evaluation results are not just about the model — they're partly about which decoding configuration the evaluator happened to use. **Worth following when:** you want to understand the generation-side decisions that affect what an LLM evaluation actually measures, beyond the model's parametric state. **Topics:** nucleus sampling and top-p decoding; the influence of decoding strategy on LLM evaluation results; text-generation theory and methodology. **Key works:** "The Curious Case of Neural Text Degeneration" introducing nucleus sampling (2019, lead author); ongoing publications on text-generation evaluation; UChicago Conceptualization Lab. --- ## Ashish Sabharwal Source: https://ansmeter.com/whom-to-read/ashish-sabharwal **Works on:** What classical formal-reasoning research can contribute to evaluating whether modern LLMs actually reason — and how to tell that from reasoning-shaped language that happens to land on the right answer. Sabharwal came to LLM research from a background in formal reasoning and satisfiability — areas where "does the system reason correctly" has a precise meaning, measured against logical specifications and verifiable by operational tests. That background reads through his more recent work on LLM reasoning evaluation: a tendency to look at whether the structure of the reasoning matches the structure the problem requires, on top of whether the final answer happens to be correct. ARC and similar benchmarks are an output of that posture — designed so that surface fluency does not substitute for actual inference. **Worth following when:** you want LLM reasoning evaluation grounded in the older formal-reasoning tradition, where "correctness" had operational meaning before deep learning arrived. **Topics:** formal reasoning and SAT-solving applied to LLM evaluation; reasoning-trace inspection in QA benchmarks; the bridge between classical AI inference and current LLM behavior. **Key works:** AI2 Reasoning Challenge / ARC (2018, co-author); pre-LLM-era work on tractable inference and SAT; recent LLM-era reasoning-evaluation publications from AI2. --- ## Benno Stein Source: https://ansmeter.com/whom-to-read/benno-stein **Works on:** Building evaluation infrastructure that turns researcher disagreement into something technically resolvable. For most of NLP and IR history, when two papers reported different numbers on the "same" benchmark, the disagreement could not be settled — different splits, different prompts, different code, different machines. Stein's Webis group built TIRA, a platform that runs submitted code in a controlled environment so that "I ran your code on this data" becomes a literally executable claim. The same infrastructure thinking drives the PAN evaluation campaigns he has co-chaired since 2009: each task includes its evaluation protocol as a binding part of the task definition. **Worth following when:** you want evaluation infrastructure that produces disagreements with technical resolution paths, not just disagreements you can publish about. **Topics:** reproducible evaluation infrastructure (TIRA); long-running shared-task design (PAN); the difference between evaluation conventions and enforceable evaluation protocols. **Key works:** TIRA Integrated Research Architecture (2007, ongoing); PAN shared-task series (2009, ongoing, co-chair); Webis group publications on retrieval and computational argumentation. --- ## Björn Schuller Source: https://ansmeter.com/whom-to-read/bjorn-schuller **Works on:** What language technology has to evaluate when the input is not text but speech — including the emotional, paralinguistic, and individual-speaker signals that text-only methods discard. Schuller has organized the INTERSPEECH Computational Paralinguistics Challenge for over a decade — annual benchmarks on emotion, sincerity, native-language, depression, and other paralinguistic categories that voice-based systems either pick up on or miss. His openSMILE feature extractor is the toolkit much of academic affective-speech research still runs on, and his published record covers the evaluation methodology side of that subfield from its modern beginnings. For Ansmeter, as AI engines move from text-only interfaces into voice-mode products, this is the literature that already worked out which paralinguistic signals matter, which can be reliably measured, and which still resist automated evaluation. **Worth following when:** you need to evaluate language-model behavior in voice mode and want methodology for measuring what the speech channel carries beyond the literal words. **Topics:** computational paralinguistics evaluation (emotion, affect, individual speaker characteristics); the INTERSPEECH challenges body of work; affective-speech feature extraction (openSMILE). **Key works:** INTERSPEECH Computational Paralinguistics Challenge organization (annual, 2009 onward); openSMILE open-source feature extractor (ongoing); affective speech-processing publications from Imperial and TU Munich. --- ## Charles L. A. Clarke Source: https://ansmeter.com/whom-to-read/charles-l-a-clarke **Works on:** What systematic, replicable evaluation of question-answering systems requires — now that the systems being evaluated are language models rather than retrievers. Clarke spent three decades inside the TREC evaluation tradition, where rigorous comparison of retrieval systems involved shared corpora, shared queries, shared relevance judgments, and an explicit protocol for resolving disagreements between assessors. His 2023 paper "Evaluating Open-Domain QA in the Era of LLMs" carries that discipline forward into the current moment, pointing out that most LLM-based QA evaluation has quietly dropped most of those guardrails — single-source ground truth, automatic judges that have not been validated against humans, no protocol for answers that are technically correct but stylistically different from the reference. The methodological reset he argues for is closer to the original TREC posture than to anything currently in vogue. **Worth following when:** you want to know what rigorous QA evaluation looked like before LLMs made everyone forget the rules, and which of those rules should come back. **Topics:** TREC-style evaluation methodology; QA evaluation in the era of LLM-generated answers; assessor disagreement and ground-truth construction. **Key works:** Information Retrieval: Implementing and Evaluating Search Engines (textbook, 2010, with Büttcher and Cormack); "Evaluating Open-Domain QA in the Era of LLMs" (2023, co-author); decades of TREC organizing and participation. --- ## Chelsea Finn Source: https://ansmeter.com/whom-to-read/chelsea-finn **Works on:** How language and learning systems adapt to a new task with very few examples — and what the theoretical structure of that adaptation tells you about what they have or haven't actually learned. Finn's MAML (Model-Agnostic Meta-Learning, 2017, lead author) formalized the question of meta-learning at architectural level: train a model not to perform a specific task, but to be quickly fine-tunable to any task in a distribution it's seen related examples from. The framework predates the in-context-learning era by several years, but its theoretical vocabulary — task distributions, adaptation steps, the loss landscape of fine-tuning — is what modern LLM in-context-learning research either explicitly builds on or independently rediscovers. For Ansmeter, the meta-learning lens applies directly: when an AI engine answers a probe query it's never seen before, the question of whether it "knows" the answer or is adapting on the fly is methodologically equivalent to the few-shot adaptation questions Finn's work formalized. **Worth following when:** you want theoretical grounding for how LLMs handle queries they weren't specifically trained for, and want the meta-learning literature that worked out the relevant frameworks before LLMs arrived. **Topics:** model-agnostic meta-learning (MAML); the theoretical structure of few-shot adaptation; the bridge between meta-learning and in-context learning in LLMs. **Key works:** MAML: Model-Agnostic Meta-Learning (2017, lead author); long body of work on meta-learning and robotic learning; Stanford and Physical Intelligence publications on RL/robotics. --- ## Chris Callison-Burch Source: https://ansmeter.com/whom-to-read/chris-callison-burch **Works on:** Whether human readers can tell apart text produced by a language model from text produced by humans — and how that distinguishability decays as models improve. Callison-Burch's "Real or Fake Text?" line of research (with Liam Dugan and others, 2020s) ran interactive experiments where readers were asked to mark the point in a text where the human author stopped and an AI continuation began. The results have tracked the progress of LLMs from an angle most evaluation skips — through 2020 the cut-off point was easy to find; by 2023 readers struggled to find it at all, even when motivated. His earlier work pioneered crowdsourcing as a primary instrument for NLP evaluation, which gives the current line additional weight: he is unusually qualified to say what human evaluators can and cannot detect under realistic conditions. **Worth following when:** you need to design human evaluation of LLM-generated text and want methodology informed by what humans actually notice versus what they think they notice. **Topics:** human distinguishability of AI vs. human text; crowdsourcing methodology for NLP evaluation; machine translation evaluation lineage from the statistical era. **Key works:** "Real or Fake Text?" line of human-vs-AI distinguishability experiments (2020 onward); crowdsourcing-for-NLP methodology papers (early 2010s); WMT-era contributions to machine-translation evaluation. --- ## Christof Monz Source: https://ansmeter.com/whom-to-read/christof-monz **Works on:** Year-over-year systematic evaluation of machine translation across dozens of language pairs — and what that tracking reveals about which translation problems are getting solved and which aren't. Most NLP evaluation is one-shot: a benchmark drops, gets saturated, gets replaced. Monz has been co-organizing the Findings of the WMT campaigns for over a decade, producing one of the few datasets in the field with longitudinal structure — the same translation task, evaluated by the same protocol, across the same language pairs, year after year. The accumulated record shows what actually transferred from the old statistical MT era to the neural era to the LLM era, and where MT progress has stalled despite the impression of universal improvement. **Worth following when:** you want evaluation methodology that's been validated across a decade of repeated annual application — instead of a benchmark that's six months old. **Topics:** annual machine translation evaluation (WMT findings); longitudinal tracking of MT progress across language pairs; the methodology lineage from statistical to neural to LLM-era MT. **Key works:** co-organization of Findings of the WMT (annual, since 2008); UvA Language Technology Lab publications on multilingual MT; ongoing work on neural MT methodology. --- ## Christopher Manning Source: https://ansmeter.com/whom-to-read/christopher-manning **Works on:** The argument that meaning, as humans use the word, is not what large language models trade in. Manning's body of work — GloVe, the foundational Stanford NLP curriculum, a generation of his PhD students who fanned out across industry labs — gives him a vantage point from which most current LLM coverage looks like a category error. His sharpest recent writing argues that when a model appears to "understand" something, we are watching statistical regularity in human language being mistaken for the conceptual structure those patterns ride on. The position is rare for someone with his standing and his institutional incentives, which is part of why it's worth taking seriously. **Worth following when:** you want to read the gap between "the model produced the right words" and "the model knew what those words meant", written by someone old enough in the field to know what got lost in the rebranding from NLP to AI. **Topics:** what distributional semantics can and can't claim about meaning; structural critiques of current LLM evaluation practice; the long arc of language technology before LLMs. **Key works:** GloVe word vectors (2014, co-lead author with Pennington and Socher); "Foundations of Statistical Natural Language Processing" (1999, with Schütze); long body of public critiques of current LLM evaluation practice. --- ## Christopher Potts Source: https://ansmeter.com/whom-to-read/christopher-potts **Works on:** What linguistic structure language models actually represent — and what they only seem to. Potts comes at language models from a linguistics-and-philosophy background, which gives him an unusual angle on what current evaluation methodology lets us conclude. Stanford Sentiment Treebank, which his group co-built more than a decade ago, is still cited in nearly every paper about sentiment classification — partly for the dataset itself, partly because it made compositional structure rather than word polarity the unit of evaluation. His more recent work on dynamic adversarial benchmarks like DynaSent argues a quieter point: a static test set goes stale the moment it appears, because subsequent models train on its echo in the corpus. **Worth following when:** you want a linguist's take on what an LLM benchmark result does and does not let you claim. **Topics:** sentiment and natural language inference as test cases for compositional meaning; adversarial and dynamic benchmark design; the linguistics behind LLM evaluation. **Key works:** Stanford Sentiment Treebank (2013, co-author); DynaSent dynamic benchmark (2021); ongoing work on compositional generalization tests for LLMs. --- ## Colin Raffel Source: https://ansmeter.com/whom-to-read/colin-raffel **Works on:** Whether a single language model can do all NLP tasks at once when they're all framed as text-in-text-out — and whether that unification holds across languages. T5 (Text-to-Text Transfer Transformer, 2020, lead author) made a methodological argument that the field had been resisting: every NLP task can be framed as the same kind of problem — give the model a string, ask for a string — and a single model trained on that framing handles them all. The follow-up T0 paper (2021) extended the argument to zero-shot multilingual generalization, showing that the same approach scales across languages with only modest performance loss. The combined line is the methodological precursor to almost every instruction-tuned multilingual LLM shipped since, including the ones Ansmeter currently evaluates. **Worth following when:** you want to understand the architectural and training-data choices that made multilingual zero-shot generalization possible — and what those choices imply for evaluation methodology. **Topics:** unified text-to-text framing of NLP tasks (T5); multilingual zero-shot generalization (T0); the lineage from research models to deployed multilingual LLMs. **Key works:** T5: Text-to-Text Transfer Transformer (2020, lead author); T0 multitask-prompted multilingual generalization (2021, co-author); ongoing publications on model memorization and dataset attribution. --- ## Dan Jurafsky Source: https://ansmeter.com/whom-to-read/dan-jurafsky **Works on:** Making natural language processing a teachable discipline — and using that perspective to read where the field's current evaluation practices fit, and don't fit, into its longer history. The textbook Jurafsky co-wrote with James Martin, Speech and Language Processing, has been the entry point for graduate NLP students for over two decades — the third edition tracks how the field absorbed deep learning while keeping the linguistic and probabilistic foundations available to anyone who wants them. His broader work runs against the grain of modern NLP in a productive way: computational sociolinguistics, the language of food, courtroom dialogue analysis, the kinds of projects that treat language as a thing humans do in social settings rather than as a benchmark to be saturated. The wider lens gives his current commentary on LLM evaluation an unusual depth — he can locate any current methodological argument in the 50-year arc of language technology. **Worth following when:** you want a senior NLP figure who treats the field as a continuous discipline. **Topics:** the standard graduate NLP curriculum (Speech and Language Processing); computational sociolinguistics and language-in-social-context analysis; the long arc of NLP as a unified field. **Key works:** Speech and Language Processing (1st ed. 2000, 3rd ed. 2023, with Martin); The Language of Food (2014); long body of work on computational sociolinguistics and social NLP. --- ## Danqi Chen Source: https://ansmeter.com/whom-to-read/danqi-chen **Works on:** Whether a language model that produces an answer with citations is actually grounding the answer in those citations — or just attaching plausible references after the fact. ALCE (2023, with Chen as senior author) is the benchmark that put the question on the table: when an LLM produces text-with-citations, do the cited sources actually support the claims they're attached to, or are the citations decorative? The framework decomposes attribution into measurable components — citation precision (the citation supports the claim), citation recall (every claim that should be cited is), and the interaction between the two — and shows that current LLMs vary enormously on which they get right. For Ansmeter, which audits whether AI engines mention brands with appropriate sourcing, this is the closest precedent for the methodology: attribution as a separate evaluation axis from raw answer accuracy. **Worth following when:** you want to evaluate whether a model's citations actually do the work they appear to be doing, or are just stylistic markers attached to plausible-sounding answers. **Topics:** automatic evaluation of LLM citation behavior (ALCE); attribution precision and recall as separate metrics; the gap between citations as decoration and citations as evidence. **Key works:** ALCE: Enabling Large Language Models to Generate Text with Citations (2023, senior author); Dense Passage Retrieval (2020, co-author with Yih); Princeton NLP publications on retrieval-augmented LMs. --- ## David Jurgens Source: https://ansmeter.com/whom-to-read/david-jurgens **Works on:** Whether language models that handle factual questions cleanly can also handle the kind of social knowledge that determines what humans actually mean when they say things. The 2023 paper "Do LLMs Understand Social Knowledge?", with Jurgens as senior co-author, ran modern LLMs through a battery of tests built from sociolinguistics research — implicature, indirect speech acts, conventionalized politeness markers, the social inferences a competent human listener makes without thinking about them. The results were mixed enough to matter: LLMs that scored high on factual QA benchmarks made systematic errors on social inferences that any reasonably socialized human gets right, especially when the social context was non-American or non-mainstream. For Ansmeter, this connects to a question we have to think about — when an engine recommends one brand instead of another, what social signal is the engine reading from the query, and would a human reading the same query interpret that signal the same way. **Worth following when:** you want LLM evaluation grounded in sociolinguistics — testing whether the model understands what humans mean, not just what they literally say. **Topics:** sociolinguistic evaluation of LLMs; computational social science with language models; the gap between factual competence and social competence in AI systems. **Key works:** "Do LLMs Understand Social Knowledge?" (2023, senior co-author); body of work on sociolinguistics and computational social science; U Michigan publications connecting NLP and social science research. --- ## Daxin Jiang Source: https://ansmeter.com/whom-to-read/daxin-jiang **Works on:** How to pre-train an encoder so that the embeddings it produces are good for retrieval — without ever supervising on a retrieval task. The standard practice for dense-retrieval models was to start from a generic encoder (BERT, RoBERTa) and fine-tune on retrieval-specific objectives — query-document pairs, contrastive losses. SimLM (2022, with Jiang as senior author) proposed a different pre-training step that produced encoders already biased toward retrieval-useful representations, so that subsequent fine-tuning needed far less labelled data to match state-of-the-art quality. The methodological consequence sits exactly where Ansmeter has to think about it: the embedding step that decides which documents the retriever sees as similar is a function of pre-training choices most papers never describe. **Worth following when:** you want to understand the pre-training-side decisions that shape what a dense retriever can and cannot find. **Topics:** retrieval-aware pre-training methodology; the gap between generic encoders and retrieval-tuned ones; what pre-training choices determine in downstream RAG behavior. **Key works:** SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval (2022, senior author); broader pre-LLM-era work on commercial-scale web NLP at Bing/Cortana. --- ## Diyi Yang Source: https://ansmeter.com/whom-to-read/diyi-yang **Works on:** How language models behave when the task is social rather than informational — persuasion, support, conflict, politeness. Yang led the 2023 "Is ChatGPT a General-Purpose NLP Solver?" study that gave the first sober task-by-task answer to a question everyone had been assuming: across more than twenty established NLP tasks, ChatGPT was strong on a few, mediocre on most, and bad on some — a pattern that complicated the narrative of broad generality. Her SALT Lab continues the harder line of inquiry: language models in roles where the right answer depends on social context — counselor, conflict mediator, persuasion target — and where the failure modes look different from what standard benchmarks reveal. **Worth following when:** you want to know how LLMs behave outside the kinds of tasks they were tuned for, especially tasks where the human stakes are higher than benchmark scores. **Topics:** task-level evaluation of LLMs on established NLP benchmarks; computational social science with language models; social context as a dimension of evaluation. **Key works:** "Is ChatGPT a General-Purpose NLP Solver?" (2023); SALT Lab publications on socially-grounded NLP (2022, ongoing). --- ## Dragan Gašević Source: https://ansmeter.com/whom-to-read/dragan-gasevic **Works on:** Whether AI tools used in educational contexts actually help the learners they're built for — and what kind of evaluation infrastructure that question requires beyond conventional ML benchmarks. Gašević is one of the founders of learning analytics as a research discipline, and the founding chair of the Learning Analytics and Knowledge (LAK) conference series. His longer career has built out the methodological infrastructure for evaluating AI-in-education at the scale of actual schools and universities — measuring not benchmark accuracy but learning outcomes, equity effects, and unintended pedagogical consequences across populations of students. For Ansmeter, the parallel is methodologically direct: a deployed AI system whose evaluation has to account for human outcomes in context. **Worth following when:** you want LLM evaluation methodology informed by educational-AI deployment experience — a domain where evaluation in real human contexts has been the price of entry for over a decade. **Topics:** learning analytics as a research discipline; AI-in-education evaluation methodology; outcome-based assessment of deployed AI systems. **Key works:** founding of Learning Analytics and Knowledge (LAK) conference series (2011 onward); body of work on learning analytics methodology; Monash and Society for Learning Analytics Research publications. --- ## Ee-Peng Lim Source: https://ansmeter.com/whom-to-read/ee-peng-lim **Works on:** Whether asking a language model to make a plan before solving a problem produces better reasoning than telling it to think step by step. The dominant zero-shot prompting trick for reasoning has been "Let's think step by step" — open-ended, no enforced structure, hope the model finds its own path. Plan-and-Solve Prompting (2023, with Lim as senior author) added a small but consequential structural change: first ask the model to articulate a plan for solving the problem, then execute that plan step by step. The empirical result was a modest but consistent improvement on math and reasoning benchmarks, and a methodological one — explicit planning leaves an artifact in the trace, so you can examine what the model thought it was about to do and where it deviated. **Worth following when:** you want prompting techniques that produce inspectable reasoning artifacts rather than just hoping the answer falls out at the end. **Topics:** plan-and-execute prompting structures; the difference between zero-shot CoT and explicit-plan prompting; inspectable reasoning artifacts in LLM evaluation. **Key works:** Plan-and-Solve Prompting (2023, senior author); long-arc body of work on data mining and social computing from his SMU group. --- ## Emma Strubell Source: https://ansmeter.com/whom-to-read/emma-strubell **Works on:** Whether the compute and energy cost of training and serving language models belongs in the headline of an evaluation, where accuracy currently sits alone. Strubell's 2019 paper "Energy and Policy Considerations for Deep Learning in NLP" put numbers to a thing the field had been quietly ignoring: training a single large language model could produce CO2 emissions comparable to several cars over their lifetimes, and the gap between research-paper accuracy reports and the carbon cost of producing them was widening with every new state-of-the-art. The work reframed efficiency from a nice-to-have engineering concern into an evaluation dimension that should show up next to accuracy in every model comparison. For Ansmeter, which evaluates models that vary by orders of magnitude in inference cost, this is the methodological backing for treating "good answer" and "good answer at what cost" as different questions. **Worth following when:** you want to report or audit an LLM result with compute and energy cost treated as first-class numbers. **Topics:** energy and compute cost of language-model training and serving; equitable NLP and access-to-compute as research problems; efficient model design as a methodology in its own right. **Key works:** "Energy and Policy Considerations for Deep Learning in NLP" (2019, lead author); Green AI position paper (2019, co-author); ongoing CMU work on efficient and equitable NLP. --- ## Emmanuel Candès Source: https://ansmeter.com/whom-to-read/emmanuel-candes **Works on:** How to put valid statistical uncertainty intervals around any model's predictions — including language models — without assuming you know what kind of error distribution to expect. Candès is one of the founding figures of conformal prediction, a statistical framework that produces calibrated prediction intervals — guaranteed coverage of the true value at a specified confidence level — without making distributional assumptions about the underlying model. The framework predates the LLM era by years and is essentially model-agnostic, which means it transfers to language models without modification: you can wrap any LLM in conformal prediction and obtain rigorous statements about the reliability of its outputs. For Ansmeter, which reports comparative scores across engines, this is the statistical foundation needed to attach actual confidence intervals to those numbers rather than relying on the field's habit of reporting point estimates as if they were facts. **Worth following when:** you need rigorous statistical uncertainty quantification for ML outputs — including LLM outputs — and want a framework that doesn't depend on the model being well-understood. **Topics:** conformal prediction and distribution-free uncertainty quantification; compressed sensing and signal processing foundations; statistical methodology for high-dimensional inference. **Key works:** foundational compressed sensing publications (mid-2000s); conformal prediction framework and applications (2008 onward, multiple key papers); Stanford statistics body of work. --- ## Eric Horvitz Source: https://ansmeter.com/whom-to-read/eric-horvitz **Works on:** The qualitative shape of language-model capability — what it looks like as a thing, and whether we have the vocabulary to describe it before we have the methodology to measure it. The Microsoft Research paper "Sparks of Artificial General Intelligence", which Horvitz co-authored in 2023, was the field's most-cited qualitative evaluation of GPT-4 — a hundred-plus pages of examples showing the model doing things its predecessors couldn't, framed as preliminary evidence of capabilities approaching general intelligence. The paper drew immediate criticism for its lack of systematic protocol, but the criticism mostly missed what made the document influential: it gave the industry a vocabulary for what it was looking at before anyone had built rigorous instruments to measure it. Horvitz's longer career — probabilistic AI, medical decision-support, AAAI presidency, the build-out of Microsoft's responsible-AI infrastructure — gives him standing the casual reader of "Sparks" usually misses. **Worth following when:** you want to read the senior-industrial perspective on what current LLMs are capable of, even when (especially when) that perspective is methodologically informal in ways that Ansmeter's own evaluation is the antidote to. **Topics:** qualitative evaluation of LLM capability (Sparks); probabilistic AI and decision-theoretic methods; the responsible-AI infrastructure inside large model labs (Aether). **Key works:** "Sparks of Artificial General Intelligence: Early experiments with GPT-4" (2023, senior author); decades of probabilistic AI and medical decision-support publications; Aether Working Group on responsible AI. --- ## Furu Wei Source: https://ansmeter.com/whom-to-read/furu-wei **Works on:** Using a language model's own parametric knowledge to make retrieval find the document the user was actually looking for. The standard direction of information flow in retrieval-augmented systems is from retrieval to generation: fetch documents first, then write the answer. Query2doc (2023, with Wei as senior author) inverts that order at the front end — given the user's query, have the language model first generate a plausible-looking answer document, then use that synthetic document as additional signal in the retrieval step. The trick works because such generated documents, while unreliable on specific facts, are usually reliable on what kind of document the user implicitly expects to find — and that's exactly the signal vector retrievers need. **Worth following when:** you want to think about retrieval and generation as bidirectional rather than as a one-way pipeline. **Topics:** query expansion via LLM-generated synthetic documents; the inverse direction of standard RAG; foundation-model contributions to retrieval methodology. **Key works:** Query2doc: Query Expansion with Large Language Models (2023, senior author); broader foundation-model and pretraining contributions including BEiT, LayoutLM family, and MiniLM. --- ## George J. Pappas Source: https://ansmeter.com/whom-to-read/george-j-pappas **Works on:** How quickly an automated attacker can find prompts that break a language model's safety alignment — and what that means for evaluating model robustness as such. Pappas came to LLM research from a long career in control theory and formal methods, where the question "how does this system fail" is treated as a primary design constraint, baked into the system from the start. His 2023 paper "Jailbreaking Black-Box LLMs in Twenty Queries" applies that posture to alignment: an attacker LLM is given the target model as a black box and instructed to find an input that bypasses the target's safety training, iteratively refining its prompts based on the target's responses. The result — that twenty queries on average suffice — reframes the conversation about LLM safety: alignment training is not a binary fix but a graded barrier whose height can be measured in attacker effort. **Worth following when:** you want to understand LLM safety as an attack-surface measurement problem informed by control theory and formal methods. **Topics:** automated jailbreaking of black-box LLMs (PAIR algorithm); safety-as-attack-surface measurement; control-theory methodology applied to LLM behavior. **Key works:** "Jailbreaking Black-Box LLMs in Twenty Queries" (2023, senior author, PAIR algorithm); long-arc body of work on control theory and formal verification of autonomous systems. --- ## Gideon Mann Source: https://ansmeter.com/whom-to-read/gideon-mann **Works on:** What language models look like when they're trained for a single high-value vertical — finance, in this case — and what evaluation that specialization requires beyond general-purpose benchmarks. Mann co-led BloombergGPT (2023), one of the first major vertical large language models — a 50-billion-parameter model trained primarily on Bloomberg's financial-document corpus, designed to outperform general-purpose LLMs on financial NLP tasks. The paper's evaluation section is methodologically interesting in its own right: side-by-side comparison with general-purpose LLMs on both financial-specific tasks (where BloombergGPT was supposed to win) and general benchmarks (where it had to remain competitive). For Ansmeter, the literature on vertical LLM evaluation matters because it answers a question we'll eventually face — how do you fairly evaluate a domain-specialist model against a generalist on queries that span both kinds of competence. **Worth following when:** you need methodology for evaluating domain-specialist LLMs against general-purpose ones, or want the literature on what vertical LLM training actually buys you. **Topics:** vertical-domain LLM development (BloombergGPT); financial-NLP evaluation; the methodological challenge of comparing specialist and generalist models on overlapping tasks. **Key works:** BloombergGPT: A Large Language Model for Finance (2023, co-author); long body of ML research at Google and Bloomberg; ongoing publications on financial NLP at scale. --- ## Graham Neubig Source: https://ansmeter.com/whom-to-read/graham-neubig **Works on:** Connecting retrieval, generation, and evaluation into a single working system — and asking when the connections actually hold. Neubig's published trajectory crosses three areas that get treated as separate fields by most other researchers: machine translation and its evaluation, retrieval-augmented generation, and the agentic LLM systems that sit on top of both. FLARE (Forward-Looking Active REtrieval, 2023, with Jamie Callan and others) sits in the middle of that triangle — a system that retrieves not just once at the start but iteratively, predicting what the next sentence needs before generating it. Reading him is a way to keep the engineering layer in view: a benchmark number means little if the system that produced it didn't actually retrieve, integrate, and verify the right things in the right order. **Worth following when:** you need to think about RAG, evaluation, and agentic systems as one stack rather than three loosely coupled research lines. **Topics:** forward-looking iterative retrieval; MT and NLG evaluation methodology; open-source agentic LLM systems. **Key works:** FLARE / Active Retrieval-Augmented Generation (2023, co-author); compare-mt and SacreBLEU-adjacent MT evaluation tooling; OpenDevin / All Hands AI agentic system work (2024, ongoing). --- ## Hannaneh Hajishirzi Source: https://ansmeter.com/whom-to-read/hannaneh-hajishirzi **Works on:** When a language model should reach for an external knowledge source — and when its own parametric memory is enough. "When Not to Trust Language Models" (2023, with collaborators across UW and AI2) is the cleanest formulation of a question the field had been talking around: a language model's internal knowledge is uneven by topic and by frequency, and an honest evaluation has to distinguish facts the model actually knows from facts it merely sounds confident about. Hajishirzi's broader work — leading the open-model line at AI2 that produced OLMo and the Tulu post-training family — extends the same logic upstream: if you want to study why models hallucinate, you have to study models you can open up, not ones served behind a closed API. **Worth following when:** you want factuality evaluation methodology that works on open-weight models you can take apart, rather than only on whatever the latest closed API returns today. **Topics:** parametric vs. retrieved knowledge in LLMs; atomic-fact factuality scoring; the open-model evaluation stack (OLMo, Tulu). **Key works:** "When Not to Trust Language Models" (2023, co-author); OLMo open language model (2024, co-author); Tulu post-training framework (2023, co-author). --- ## Igor Mordatch Source: https://ansmeter.com/whom-to-read/igor-mordatch **Works on:** What it means to evaluate a language model that takes consequential actions through its text — not only producing answers but operating in environments that respond. Mordatch's research line — embodied reinforcement learning, multi-agent emergent communication, language-conditioned agents — has been at the intersection of LLMs and the environments they could potentially act on. As LLMs evolve into agent systems that take consequential actions (browse, execute code, talk to APIs, manipulate documents), the evaluation question changes shape: the right thing to measure is what happened in the world as a result of the action sequence, beyond the surface correctness of the text produced. His earlier OpenAI work on multi-agent emergent language and current DeepMind work on language-conditioned action policies sit at the methodological frontier for agent-grade evaluation. For Ansmeter, the agent transition is on the horizon — when AI engines move from answering brand-questions to taking actions on behalf of users, the evaluation methodology has to evolve to match. **Worth following when:** you want to understand evaluation methodology for AI-agent systems — where the unit of analysis is the action sequence and its consequences, not the text alone. **Topics:** language-conditioned reinforcement learning agents; multi-agent emergent communication; the methodological transition from text-output evaluation to action-sequence evaluation. **Key works:** body of work on multi-agent emergent communication and language-conditioned RL (OpenAI 2017–2020, then Google DeepMind); embodied RL publications; ongoing language-agent research at DeepMind. --- ## Ion Stoica Source: https://ansmeter.com/whom-to-read/ion-stoica **Works on:** The distributed-compute infrastructure that almost every modern LLM is either trained on, served from, or evaluated through. Stoica's name is on the founding papers of Spark (distributed data processing), Ray (distributed Python for ML workloads), and the broader stack vLLM uses to serve language models efficiently — artifacts that have become the de facto substrate for ML in industry. The pattern is consistent: design a system around the actual computational bottlenecks researchers and engineers face, ship it open-source with a commercial company alongside, and let the ecosystem prove the design choices. For Ansmeter, which runs hundreds of audits against multiple LLM engines, the infrastructure side Stoica's group built is not background — it's the layer that determines how feasibly you can run any large-scale model comparison at all. **Worth following when:** you want to understand the infrastructure constraints that shape what large-scale LLM evaluation can actually do, beyond the methodological framing. **Topics:** distributed ML systems and infrastructure (Spark, Ray, vLLM); the economics and engineering of large-scale model serving; the open-source-plus-commercial-company pattern in ML infrastructure. **Key works:** Apache Spark (2010, co-creator); Ray distributed framework (2018, co-lead); contributions to the vLLM efficient-serving stack (2023, ongoing). --- ## Iryna Gurevych Source: https://ansmeter.com/whom-to-read/iryna-gurevych **Works on:** Building the encoder infrastructure that makes "find sentences similar to this query" a fast and reliable operation across languages and tasks. Sentence-BERT (2019, with Nils Reimers as lead author from her UKP Lab) made dense sentence encoding production-grade — fine-tuning a transformer so that two sentences with similar meaning produce nearby vectors, and packaging the result so that any developer could use it without retraining. The downstream effect is that the sentence-transformers library her group seeded became the substrate for most modern semantic-search systems, including the retrieval layers of RAG pipelines Ansmeter cares about. Her broader research program at UKP continues along the same line: practical, reusable NLP components, often with explicit evaluation methodology attached. **Worth following when:** you want to understand the sentence-encoder infrastructure that most modern retrieval-augmented systems depend on, and the methodology behind making it work. **Topics:** sentence embeddings and dense retrieval (Sentence-BERT, sentence-transformers); argument mining and computational social science applications; reusable NLP components with attached evaluation methodology. **Key works:** Sentence-BERT (2019, senior author with Reimers); the sentence-transformers library ecosystem (2020 onward); UKP Lab body of work on argument mining and applied NLP. --- ## James Zou Source: https://ansmeter.com/whom-to-read/james-zou **Works on:** What evaluation looks like when the AI being evaluated has to clear regulatory bars before deployment — and what general LLM evaluation should learn from a field that has been doing this for years. Zou's research at Stanford has tracked the FDA's growing roster of approved AI/ML medical devices and the evaluation methodology each one had to satisfy: not benchmark scores, but prospective clinical performance studies with predefined endpoints. His work also includes systematic comparisons of LLMs and clinicians on medical-question-answering, which gives the field a calibrated reference for what "this LLM is medically reliable" actually requires. For Ansmeter, biomedical AI is the methodological exemplar — a domain where the evaluation question is structurally similar to ours (a high-stakes recommendation in a context that affects real outcomes) but where the stakes have forced a longer methodological discipline than general LLM evaluation has yet developed. **Worth following when:** you want LLM evaluation methodology informed by a domain where stakes have already forced rigorous standards — and want to see what those standards actually require. **Topics:** biomedical AI evaluation; FDA-track AI/ML medical device approval methodology; the bridge between regulatory-grade evaluation and general LLM benchmarking. **Key works:** body of work on FDA-approved AI medical devices and their evaluation criteria (Stanford, 2020 onward); systematic LLM-vs-clinician comparison studies; ongoing publications on AI-in-healthcare evaluation methodology. --- ## Jamie Callan Source: https://ansmeter.com/whom-to-read/jamie-callan **Works on:** What retrieval-augmented generation looks like when you actually know the thirty-year history of information retrieval it's reinventing. Callan has been building search engines and retrieval evaluation corpora since before "RAG" existed as a term — Indri, the ClueWeb09 through ClueWeb22 datasets, the kinds of artifacts that NLP papers cite as evaluation infrastructure without realizing they were the original IR research. His more recent contribution as co-author on FLARE (Forward-Looking Active RAG, 2023) is the bridge: an IR foundational figure helping a younger generation of NLP researchers recognize which retrieval problems are genuinely new and which were solved twenty years ago. Reading him is a way to avoid the standard NLP-side mistake of treating the retrieval layer as a black box that "just does search". **Worth following when:** you want to understand the IR foundations underneath modern RAG systems — and the failure modes those foundations already mapped. **Topics:** information retrieval foundations and history; web-scale retrieval corpora (ClueWeb family); active retrieval as the bridge between IR and LLM-driven NLP. **Key works:** Indri search engine (early 2000s, ongoing); ClueWeb corpora (2009 through 2022); FLARE / Active Retrieval-Augmented Generation (2023, co-author). --- ## Jared Kaplan Source: https://ansmeter.com/whom-to-read/jared-kaplan **Works on:** Whether language-model loss is a smooth function of compute, data, and model size — and what that smoothness lets you predict (and not predict) about capabilities at larger scales. Kaplan led "Scaling Laws for Neural Language Models" (2020, lead author with OpenAI collaborators), the paper that formalized what the field had been seeing empirically: LLM training loss scales as a clean power law of compute, parameters, and dataset size, with predictable exponents. The result became the planning instrument for every subsequent training run at scale — if you know the loss curve, you can budget compute and data to hit a target performance with low surprise. The methodological limitation, which the paper itself was careful about and the field has since rediscovered the hard way, is that smooth scaling of loss does not translate cleanly into smooth scaling of any particular capability — emergent-capability discontinuities live in the gap between aggregate loss and task-level evaluation. **Worth following when:** you want to understand what the underlying scaling theory predicts, what it doesn't, and where the methodological boundary between loss-level prediction and capability-level evaluation actually sits. **Topics:** scaling laws for language model training; the gap between aggregate-loss scaling and task-level capability prediction; the methodological foundations of large-scale LLM training. **Key works:** "Scaling Laws for Neural Language Models" (2020, lead author); subsequent scaling and capability research at Anthropic; theoretical-physics-to-ML methodology body of work. --- ## Jennifer Wortman Vaughan Source: https://ansmeter.com/whom-to-read/jennifer-wortman-vaughan **Works on:** How people actually understand and act on language-model evaluation results — and where the gap between what was measured and what gets believed. Vaughan's research has long focused on a question that most ML evaluation skips: even when you have a clean number — accuracy, calibration, fairness measure — what happens when that number meets a human decision-maker who has to use it? Her studies on interpretability find regularly that explanations meant to build user trust can produce overconfidence instead, with people accepting model outputs they should have questioned, because the explanation looked authoritative regardless of its substance. For Ansmeter, which produces evaluation reports customers will use to make six-figure decisions, the lesson is direct: the artifact we ship is not the score itself — it's whatever the customer ends up believing the score means. **Worth following when:** you produce evaluation reports for non-ML audiences and want to know how those reports actually get read. **Topics:** human interpretation of ML evaluation results; calibration and trust in human-AI collaboration; transparency frameworks for ML systems (Aether). **Key works:** "Manipulating and Measuring Model Interpretability" (2021, co-author); long-standing work on fairness and accountability in ML (FATE); WiML community organization (co-founder, 2006 onward). --- ## Jesse Dodge Source: https://ansmeter.com/whom-to-read/jesse-dodge **Works on:** What has to be in a paper about a language model for another lab to be able to verify the result. Before Dodge's work, most LLM-efficiency publications compared results obtained under incompatible conditions — different hyperparameter-search budgets, different dataset variants, undocumented normalization steps that made head-to-head numbers misleading. The "Show Your Work" reporting standard he formulated and pushed at NeurIPS and ACL as a reviewer requirement turned reproducibility from a convention into an enforceable line. The same instinct drives his Green AI work — report not just accuracy but compute cost, because "percent correct" is not a comparable number across labs without it. **Worth following when:** you want to know what would actually have to be true of an LLM publication for the result it reports to be reproducible by someone else. **Topics:** reporting standards for ML research (Show Your Work); Green AI methodology and energy-cost accounting; what open-model evaluation infrastructure looks like (OLMo, Dolma, DataDecide). **Key works:** "Show Your Work: Improved Reporting of Experimental Results" (2019); "Green AI" position paper (2019, co-author); OLMo and Dolma open-model evaluation tooling (2024). --- ## Ji-Rong Wen Source: https://ansmeter.com/whom-to-read/ji-rong-wen **Works on:** Organizing the LLM literature into something other Chinese-language NLP researchers can actually navigate from inside the academic ecosystem. Wen co-led "A Survey of Large Language Models" — at the time of publication, the most-cited overview of the field — with collaborators across Renmin University and other Chinese institutions. The framing of the survey was important: not a Western-centric overview translated for Chinese readers, but a synthesis written from the Chinese academic NLP community's vantage point, with attention to architectural decisions and training-data choices specific to multilingual and Chinese-language pretraining. His longer arc — IR foundations, MSRA tenure, current institutional roles at BAAI and Renmin — gives him standing to write that synthesis with authority on both sides of the LLM ecosystem. **Worth following when:** you want a foundational LLM-landscape survey written from inside the Chinese academic AI community rather than translated into it. **Topics:** comprehensive surveys of the LLM landscape; the Chinese-language AI research ecosystem; IR-side foundations of modern language models. **Key works:** "A Survey of Large Language Models" (2023, senior author with Zhao and collaborators); foundational IR work from MSRA and Renmin; ongoing leadership of BAAI research direction. --- ## Jian-Guang Lou Source: https://ansmeter.com/whom-to-read/jian-guang-lou **Works on:** Whether a language model that produces code actually produces code that does what the request asked for — and what evaluation looks like when correctness has a definite answer for once. Lou's MSR Asia research line has worked on the rare LLM-evaluation problem where correctness is unambiguous: code generation. Either the generated program compiles, runs, and produces the right output, or it doesn't — which means the evaluation methodology can sidestep most of the LLM-as-judge problems that plague open-ended generation evaluation. His group's contributions to in-context-learning evaluation and program-synthesis benchmarks built on this advantage: a literature that knows what "correct" means and uses that as leverage to interrogate other parts of model behavior. **Worth following when:** you want to study LLM evaluation methodology in the rare setting where the ground truth is unambiguous — and to see what that clarity buys you that open-ended evaluation lacks. **Topics:** program-synthesis evaluation methodology; in-context-learning rigor for code generation; the contrast between unambiguous-ground-truth and open-ended evaluation. **Key works:** body of work on program synthesis and code generation evaluation at Microsoft Research Asia (2018 onward); in-context-learning evaluation publications; ACM Distinguished Member contributions to applied LLM research. --- ## Jian-Yun Nie Source: https://ansmeter.com/whom-to-read/jian-yun-nie **Works on:** How search behaves differently when the query is a conversation in a non-English language — and how language models are changing that picture. Nie spent most of his career on cross-lingual and multilingual information retrieval — work that was a niche specialty when the field was largely English-centric, and that turned out to be directly relevant once LLM-based search products started shipping to non-English markets. His ConvGQR research on conversational generative query reformulation extends the same thread: when a multi-turn dialogue happens in a less-resourced language, the rewriting step that makes retrieval possible has to do more work, and many of the established assumptions break down. For Ansmeter, which audits AI engines in five languages, this is the literature that maps which assumptions in standard IR don't survive the language switch. **Worth following when:** you want IR research that takes non-English search behavior seriously rather than treating it as a porting problem from English. **Topics:** cross-lingual and multilingual information retrieval; conversational generative query reformulation (ConvGQR); the language-specific failure modes of search systems. **Key works:** foundational work on cross-lingual IR through the 1990s and 2000s; ConvGQR conversational query reformulation (2023, senior author); ongoing RALI lab publications on multilingual search. --- ## Jianfeng Gao Source: https://ansmeter.com/whom-to-read/jianfeng-gao **Works on:** What large language models are as a class of systems — taxonomically, architecturally, and in terms of what they actually inherit from the longer history of neural NLP. Gao led the 2024 "Large Language Models: A Survey" — one of the broader synthesis papers of the LLM era, organizing the literature not by paper but by the questions that distinguish one architectural lineage from another (decoder-only versus encoder-decoder, instruction-tuning lineages, alignment-training variations, modality extensions). The framing is unusual in that it treats LLMs as a research artifact to be classified and analyzed at the species level, not as a single thing to be cheered for or against. His longer body of work in conversational AI and grounding gives that taxonomical view depth — he has seen which architectural choices survive contact with deployment and which don't. **Worth following when:** you want a taxonomically organized map of the LLM landscape from someone with the industrial-research vantage point to know which distinctions actually matter in practice. **Topics:** large-scale survey of the LLM landscape; conversational AI and grounding; the architectural lineages of modern language models. **Key works:** "Large Language Models: A Survey" (2024, senior author with collaborators); long body of work on conversational AI and language grounding; Microsoft Research Redmond publications on foundation models. --- ## Jie Zhou Source: https://ansmeter.com/whom-to-read/jie-zhou **Works on:** Whether language models can be trusted to evaluate other language models in production NLP pipelines. The 2023 "Is ChatGPT a Good NLG Evaluator?" preliminary study, with Zhou as senior author, was one of the first papers to seriously ask the question the field had been treating as a foregone conclusion — can a single closed-source LLM be the scoring oracle for an entire research line, and what happens when you actually check? The findings were mixed: ChatGPT-as-judge correlated well with humans on some NLG tasks and badly on others, with biases that didn't show up until you stratified by task type. Zhou's broader work is in industrial conversational AI, which lends a particular angle to the question: when your dialogue product talks to hundreds of millions of users, "the evaluator is mostly right" is not a tolerable answer for the cases where it isn't. **Worth following when:** you want a methodologically cautious early treatment of LLM-as-judge, written by someone who has to make scoring decisions count in production. **Topics:** early empirical studies of LLM-as-judge reliability; task-stratified evaluator bias; industrial conversational AI evaluation. **Key works:** "Is ChatGPT a Good NLG Evaluator? A Preliminary Study" (2023, senior author); ongoing publications on large-scale conversational systems. --- ## Jimmy Lin Source: https://ansmeter.com/whom-to-read/jimmy-lin **Works on:** Making information retrieval reproducible enough that an LLM researcher and an IR researcher can run the same experiment and get the same answer. Lin's Anserini and Pyserini toolkits became the default reproducibility layer for retrieval research the moment they appeared — when an NLP paper now reports a BM25 baseline number, that number usually comes from his code. The same posture carries into HyDE (Hypothetical Document Embeddings, 2023, senior author): the idea was to use a language model to generate a synthetic answer document, embed it, and retrieve against that embedding — a method clean enough that other groups could replicate it without negotiating about implementation details. Reading Lin is the way to understand which retrieval results in current LLM papers actually mean something and which depend on undocumented baselines. **Worth following when:** you want to know whether a retrieval-related result in a paper is reproducible — and how to set things up so that yours will be. **Topics:** dense retrieval and zero-shot retrieval via LLM-generated embeddings; reproducibility infrastructure for IR research; the bridge between IR and neural NLP communities. **Key works:** Anserini IR toolkit (2017, lead); Pyserini (2021, lead); HyDE: Precise Zero-Shot Dense Retrieval (2023, senior author). --- ## Jimmy Xiangji Huang Source: https://ansmeter.com/whom-to-read/jimmy-xiangji-huang **Works on:** How information retrieval scales over the messy operational data that real organizations hold, as opposed to the clean benchmark corpora the field actually publishes on. Most academic IR research runs on web-scale or news-scale collections — corpora where the retrieval problem is well-defined and the documents are reasonably clean. Huang's IR&KM Lab at York has spent two decades on a more representative case: retrieval over the kind of mixed, heterogeneous, sometimes structured data that actual enterprises hold, where the failure modes look different from anything a benchmark captures. His recent LLM-evaluation work extends the same lens — how do retrieval-augmented language models behave when the corpus they retrieve from is the kind of repository an actual organization runs on, not a curated research collection. **Worth following when:** you want IR research that takes seriously the gap between benchmark corpora and the kind of data RAG systems actually encounter once deployed. **Topics:** information retrieval over heterogeneous and enterprise-scale data; the corpus-side conditions of RAG behavior; systematic evaluation of LLM-IR hybrids on non-web data. **Key works:** body of work on IR for big-data and enterprise corpora (2000s onward, York IR&KM Lab); ongoing systematic LLM-IR evaluation publications. --- ## Jindong Wang Source: https://ansmeter.com/whom-to-read/jindong-wang **Works on:** Standardizing what "evaluating an LLM" even means as a research procedure. The literature on LLM evaluation grew faster than the methodology to make those papers comparable — different prompts, different scoring, different definitions of "success" produced contradictory claims about the same models. Wang led the most-cited synthesis of that landscape (the 2023 "Survey on Evaluation of LLMs", with collaborators) and built PromptBench, an open framework for stress-testing model robustness across prompt variations — the part of the field where the gap between "the model works" and "the model works on the exact words we tested" lives. **Worth following when:** you want to design an LLM evaluation from scratch and want to know what conventions already exist before you reinvent them. **Topics:** systematic synthesis of LLM evaluation research; prompt robustness and adversarial prompts; the gap between benchmark scores and deployed behavior. **Key works:** "A Survey on Evaluation of LLMs" (2023, lead author); PromptBench framework (2023). --- ## Jochen Wirtz Source: https://ansmeter.com/whom-to-read/jochen-wirtz **Works on:** What happens when customer-facing service AI replaces, augments, or competes with human service workers — and what kinds of evaluation that change actually requires. Wirtz has spent his career on service marketing and, more recently, on service AI: the research line that asks what the deployment of chatbots, robots, and AI agents into customer-facing roles does to customer experience, trust, and brand perception. The discipline he comes from cares about evaluation methodology at the level customers actually experience — survey instruments validated across cultural contexts, longitudinal effects on brand loyalty, the difference between satisfaction at first interaction and satisfaction after the system fails. For Ansmeter, which evaluates how AI engines shape what users hear about brands, service-AI literature is the closest precedent for measuring the customer-side of that exposure. **Worth following when:** you need methodology for evaluating consumer-facing AI deployments and want the service-marketing research tradition's instruments for measuring brand and customer-experience effects. **Topics:** service AI and service marketing research; customer experience measurement methodology; the longitudinal effects of AI deployment on brand perception. **Key works:** body of work on service AI and service marketing (NUS, ongoing); co-editorship of major service-marketing journals; foundational textbook contributions to service-marketing methodology. --- ## Jonathan Berant Source: https://ansmeter.com/whom-to-read/jonathan-berant **Works on:** Whether a language model is actually doing the reasoning steps needed to answer a question — or just producing an answer that happens to be right. For complex questions — "what's the total population of countries whose capital starts with 'B'?" — the difference between knowing the answer and reasoning to it matters, but most QA evaluation collapses both into a single correctness score. Berant's BREAK and QDMR work formalized the alternative: decompose each complex question into the atomic reasoning steps it implies, then check whether the model went through those steps, not just whether the final answer matched. The DROP reading-comprehension benchmark (2019, co-author) put numbers on how much current LLMs lose when the question demands actually composing intermediate facts to reach the answer. **Worth following when:** you want to evaluate not just whether an LLM got the answer, but whether it reasoned its way there in a structurally honest way. **Topics:** question decomposition into atomic reasoning steps (QDMR, BREAK); compositional reading comprehension (DROP); the gap between answer correctness and reasoning correctness. **Key works:** BREAK dataset with QDMR annotations (2020, lead author); DROP reading-comprehension benchmark (2019, co-author); ongoing Tel Aviv and Google DeepMind work on compositional reasoning. --- ## Juanzi Li Source: https://ansmeter.com/whom-to-read/juanzi-li **Works on:** Combining structured knowledge from knowledge graphs with the statistical patterns language models learn from text — and what that combination buys you for evaluation. Li's KEG Lab at Tsinghua has been the methodological home for one of the longest-running lines on knowledge-graph-augmented language models — OpenKE for KG embeddings, KEPLER for combining KG and language pretraining, and a research thread that led, via Zhipu, to the ChatGLM open LLM line. The throughline is a commitment to language models that have an external structured-knowledge anchor as well as their parametric memory, which makes the resulting systems easier to audit — you can ask the model and you can ask the KG separately, then compare. For Ansmeter, which audits brand-mention behavior, the KG-augmented evaluation tradition is the closest precedent for separating "what the model says" from "what its sources support". **Worth following when:** you want LLM evaluation methodology that takes external structured-knowledge anchors seriously — both in model design and in audit protocol. **Topics:** knowledge-graph-augmented language models (KEPLER, OpenKE); the lineage from KG research to open Chinese LLMs (ChatGLM); auditability through external knowledge anchors. **Key works:** OpenKE knowledge-graph embedding library (2018 onward, key contributor); KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation (2021, co-author); ChatGLM open LLM line via Zhipu (2023 onward, co-founder). --- ## Julian McAuley Source: https://ansmeter.com/whom-to-read/julian-mcauley **Works on:** Whether language models trained on the open web are already doing recommendation — and what that implies for products that compete with traditional recommenders. The standard recommender system was a specialized model trained on a particular catalog and user-behavior history. McAuley's 2023 work on "Large Language Models as Zero-Shot Conversational Recommenders" demonstrated that an off-the-shelf LLM, with no recommender-system training at all, performs competitively on conversational-recommendation benchmarks — partly because the training data contained a billion implicit recommendations from forums, reviews, and listicles. For Ansmeter, this is the precise mechanism behind "the model mentions brand X and not brand Y in answer to an open-ended question" — it's conversational recommendation in everything but the label. **Worth following when:** you want to understand the recommendation behavior buried inside open-ended LLM responses, and why it isn't always called by that name. **Topics:** LLMs as zero-shot conversational recommenders; recommendation behavior in open-ended LLM outputs; the training-data conditions that made out-of-the-box LLM recommendation work. **Key works:** "Large Language Models as Zero-Shot Conversational Recommenders" (2023, senior author); Amazon product-review datasets (mid-2010s onward); foundational publications connecting NLP and recommendation. --- ## Junichi Yamagishi Source: https://ansmeter.com/whom-to-read/junichi-yamagishi **Works on:** Whether human listeners — or automated detectors — can tell AI-generated speech apart from real human speech, and how that distinguishability changes as speech synthesis improves. Yamagishi has been the primary organizer of ASVspoof, the international challenge that tracks how well automated systems can detect spoofed or synthetic speech, since the first edition in 2015. The accumulated record from the challenge series is one of the few longitudinal datasets in AI-content-detection that covers an entire technology generation — from early concatenative synthesis to current neural voice cloning. He co-organizes the parallel VoiceMOS challenge for evaluating synthetic-speech quality, which makes him uniquely positioned on both sides of the question: how good has synthesis gotten, and how detectable does that goodness remain. For Ansmeter, as AI engines move toward voice output, this is the methodology stack for evaluating what it sounds like when an LLM speaks. **Worth following when:** you need to evaluate AI-generated speech — either its quality or its detectability — and want longitudinal benchmark methodology that has tracked the technology through multiple generations. **Topics:** anti-spoofing for voice (ASVspoof); synthetic-speech quality evaluation (VoiceMOS); the longitudinal tracking of speech-synthesis progress through annual challenges. **Key works:** ASVspoof challenge series organization (2015 onward); VoiceMOS challenge co-organization; long body of work on speech synthesis at NII and Edinburgh CSTR. --- ## Jure Leskovec Source: https://ansmeter.com/whom-to-read/jure-leskovec **Works on:** What graph structure adds to the kinds of reasoning and retrieval problems language models currently handle without it — and what gets missed when relationships in data are flattened into text. Leskovec's node2vec (2016) was one of the first methods to make graph-structured data work as input to standard ML pipelines — embed nodes as vectors that preserve neighborhood information, then use the embeddings as features anywhere a vector goes. The Stanford Network Analysis Project (SNAP) datasets and methods that followed became the default substrate for graph-ML research. For Ansmeter, where the question of "how does a model know what brand to mention" sits in graph-structured-reasoning territory — which entities are connected to which contexts in training data — the graph perspective is a usefully orthogonal angle on what current LLM evaluation tends to measure only on the surface. **Worth following when:** you want to think about LLM knowledge and reasoning behavior in terms of graph structure underneath the text, rather than as a property of the text itself. **Topics:** graph machine learning and node embeddings (node2vec); the Stanford Network Analysis Project (SNAP); graph-structured reasoning as a lens on LLM behavior. **Key works:** node2vec: Scalable Feature Learning for Networks (2016, co-author); Stanford Network Analysis Project datasets and tools (2007 onward); graph-ML textbook and curriculum contributions. --- ## Kevin Chen-Chuan Chang Source: https://ansmeter.com/whom-to-read/kevin-chen-chuan-chang **Works on:** Organizing the rapidly growing literature on language-model reasoning into something a researcher new to the area can actually navigate. By mid-2023 the number of "LLMs can reason" papers exceeded any individual reader's capacity to keep up — chain-of-thought, decomposition, planning, self-consistency, in-context learning, each with multiple variants and competing claims. Chang's "Towards Reasoning in LLMs" survey (2023, with Jie Huang) gave the field its first reasonably complete map of the territory, organizing techniques by mechanism, evaluation paradigm, and empirical findings about what does and doesn't transfer across tasks. Reading the survey is the cleanest way to acquire shared vocabulary with the literature without having to re-derive it from a hundred individual paper abstracts. **Worth following when:** you need to onboard quickly to LLM reasoning research and want a structured map that's been peer-reviewed rather than a curated reading list. **Topics:** survey methodology for fast-moving subfields; taxonomy of LLM reasoning techniques; the structure of the LLM reasoning literature. **Key works:** "Towards Reasoning in Large Language Models: A Survey" (2023, with Huang); long-arc body of work on data mining, deep-web search, and NLP+IR from UIUC. --- ## Kyle Lo Source: https://ansmeter.com/whom-to-read/kyle-lo **Works on:** What's actually in a training corpus once you sit down and look at it document by document. Most discussion of LLM training data happens at the headline level — "trained on the web", "trained on Common Crawl". Lo's work on Dolma, the corpus behind OLMo, is what that conversation looks like when someone actually opens the file: documented decisions about which CommonCrawl snapshots, which deduplication thresholds, which filtering heuristics, which document types, all written up at the level of detail a commercial lab would call confidential. SciBERT, earlier in his career, did the equivalent move for scientific text — bake in domain-specific tokenization and pretraining from documented corpora, then ship the artifact so others can replicate. **Worth following when:** you want to understand what's inside a training corpus at the granularity needed to predict downstream model behavior — rather than at the marketing-summary level. **Topics:** open training corpora and their documentation (Dolma); domain-specific pretraining (SciBERT); the documentation standards that distinguish open from closed model development. **Key works:** Dolma open training corpus (2024, lead); OLMo open language model (2024, co-lead); SciBERT (2019, lead). --- ## Luke Zettlemoyer Source: https://ansmeter.com/whom-to-read/luke-zettlemoyer **Works on:** Whether open-weight language models can be built at frontier scale and whether their factuality can be measured at fine resolution. Zettlemoyer's name appears on a roster of work that is closer to a research program than a list of papers — open large language models that anyone can examine (OPT in 2022 and the open-model line that followed), and the evaluation infrastructure to say whether those models keep their factual claims straight. FActScore (2023, lead author Sewon Min) is the methodologically careful version: break a long generated paragraph into atomic factual claims and score each against a knowledge source, rather than score "the whole paragraph" with a single judgment. For Ansmeter the resolution matters — a model that gets one fact right and two wrong in the same sentence should not get a passing grade. **Worth following when:** you want fine-grained factuality scoring rather than a single yes-or-no judgment per response. **Topics:** atomic factuality scoring; open large language models and their evaluation; semantic parsing as a precursor to today's structured-output methods. **Key works:** FActScore (2023, co-author); OPT open language model (2022, co-author); foundational semantic parsing work in the 2000s. --- ## Maarten de Rijke Source: https://ansmeter.com/whom-to-read/maarten-de-rijke **Works on:** Whether information retrieval should be done by a system that "writes the document ID" instead of by one that searches a vector index. De Rijke has been one of the architects of generative retrieval — a methodological alternative to the standard "embed the query, search a vector index" paradigm, in which a sequence-to-sequence model is trained to directly emit the identifier of the relevant document, with no separate retrieval step. The question this opens is whether the retriever and the generator should ever have been separate components in the first place, or whether the whole pipeline was an artifact of how the field happened to evolve. For Ansmeter, the consequence is structural: if generative retrieval keeps getting better, the assumption that "retrieval" and "generation" are distinguishable layers — which underlies most current evaluation — may not hold for the next generation of engines. **Worth following when:** you want to track the line of research that questions whether retrieval and generation should be architecturally distinct at all. **Topics:** generative retrieval (Differentiable Search Index family); counterfactual evaluation in IR; conversational search systems. **Key works:** generative-retrieval line of papers (2022 onward, senior author across multiple); contributions to Foundations and Trends in Information Retrieval; ongoing UvA IRLab publications. --- ## Maarten Sap Source: https://ansmeter.com/whom-to-read/maarten-sap **Works on:** Whether language models can produce or reason about social knowledge with the same competence they show on factual tasks — and what's at stake when they can't. Sap's earlier work on COMET (Commonsense Transformers, 2019) built one of the field's first attempts at making machine-readable representations of everyday social inferences — that someone who borrowed money will probably want to repay it, that an apology implies prior wrongdoing. The same lens runs through his LLM-era work: what kinds of social and commonsense knowledge are LLMs producing fluently versus mimicking superficially, and where do the failures cluster. RealToxicityPrompts and adjacent benchmarks made that question quantitative for safety-relevant cases — models trained on the open web inherit the toxic patterns of that web in measurable ways, even when their alignment training tries to paper over the inheritance. **Worth following when:** you want to evaluate LLM behavior on social and commonsense reasoning, not just factual QA, with benchmarks that distinguish surface compliance from underlying capability. **Topics:** commonsense knowledge representation in neural models (COMET); LLM behavior on social and moral reasoning; toxicity benchmarks (RealToxicityPrompts) and what they actually measure. **Key works:** COMET commonsense transformers (2019, co-lead author); RealToxicityPrompts (2020, lead author); ongoing CMU and AI2 publications on social and safety NLP. --- ## Maosong Sun Source: https://ansmeter.com/whom-to-read/maosong-sun **Works on:** What the open-LLM ecosystem looks like when it grows out of Chinese academic NLP rather than out of Western non-profits like AI2. Sun runs the THUNLP lab at Tsinghua, the academic anchor for the OpenBMB project — a Chinese-side parallel to the AI2 OLMo effort, releasing open large models (CPM and others) along with the training infrastructure (BMTrain) needed to reproduce them. His broader work on parameter-efficient fine-tuning came out of the same instinct: take expensive techniques developed at scale by closed labs, find the small-data version that works on academic compute, and ship the result open-source. For Ansmeter, which audits AI engines across language regions, reading Sun is the way to see what the open-model alternative looks like when it grows from Chinese institutional roots rather than Western ones. **Worth following when:** you want to understand the Chinese-side open-LLM ecosystem and the parameter-efficient methods that made academic-scale LLM research possible. **Topics:** open large-model platforms from China (OpenBMB, CPM family); parameter-efficient fine-tuning methodology; THUNLP body of work on Chinese-language NLP. **Key works:** OpenBMB project body of work (2022 onward, key contributor); parameter-efficient fine-tuning publications; THUNLP foundational publications on Chinese NLP. --- ## Marco Baroni Source: https://ansmeter.com/whom-to-read/marco-baroni **Works on:** Whether neural language models can compose what they've learned into new combinations they've never seen — or whether they're really only doing sophisticated interpolation within their training distribution. Baroni's group built SCAN (2018) as a controlled test of compositional generalization: a small synthetic language where the model has to combine verbs and modifiers in patterns absent from training. Neural sequence-to-sequence models, even at the time the dominant architecture, failed spectacularly — generalizing one short held-out combination but breaking on slightly longer ones. The result generalized: subsequent papers have shown the same compositional-generalization failures in modern LLMs at scale, which means the question Baroni put on the table in 2018 hasn't been answered by scale alone. **Worth following when:** you need to assess whether a model's reasoning is genuinely compositional or merely a sophisticated form of in-distribution pattern matching. **Topics:** compositional-generalization benchmarks (SCAN and successors); distributional semantics from a linguistic angle; what LLM scale does and doesn't fix about fundamental generalization gaps. **Key works:** SCAN compositional-generalization benchmark (2018, senior author); long-arc work on distributional semantics from UPF Barcelona; ACL test-of-time award publications. --- ## Mari Ostendorf Source: https://ansmeter.com/whom-to-read/mari-ostendorf **Works on:** What language technology looks like when you've been responsible for it as deployable engineering for thirty years before LLMs arrived to redo the field. Ostendorf spent decades on the speech side of language technology — speech recognition, prosody, dialogue systems — in a tradition that took engineering-grade reliability for granted because the systems had to run on real audio for real users. That posture is increasingly relevant to LLM evaluation: as language models move from text-only demos into voice-enabled products, the engineering questions speech researchers solved for thirty years matter again. Her current work, combined with the systems-engineering tradition she comes from, makes her a useful read for anyone trying to evaluate language tech as something with deployment consequences. **Worth following when:** you want a senior systems-engineering perspective on what reliable language-technology deployment looks like — from someone whose subfield made that the price of entry. **Topics:** speech and spoken-language systems engineering; the lineage from automatic speech recognition to multimodal LLM deployment; reliability standards across language-tech subfields. **Key works:** decades of speech-recognition and prosody publications (1990s onward); UW SSLI lab body of work on dialogue and conversational systems; engineering-academy-level perspective on language-tech maturity. --- ## Mark Gales Source: https://ansmeter.com/whom-to-read/mark-gales **Works on:** Detecting hallucinations in a language model without any access to its weights or to ground truth. Most hallucination-detection methods rely on at least one of three things: the model's internal probabilities, a labeled reference answer, or a more powerful second model acting as judge. SelfCheckGPT, which Gales's Cambridge group introduced in 2023, removed all three — the idea is to ask the model the same thing several times and watch whether the answers stay self-consistent. Inconsistency across samples is read as a signal of fabrication; consistency, as a signal of grounded retrieval from training. **Worth following when:** you need to audit a commercial LLM you don't own and can't pair with a second model. **Topics:** zero-resource hallucination detection; self-consistency as an evaluation signal; the speech-modeling tradition behind much of modern LLM uncertainty work. **Key works:** SelfCheckGPT (2023); decades of speech foundation work feeding into the uncertainty methodology. --- ## Martin Potthast Source: https://ansmeter.com/whom-to-read/martin-potthast **Works on:** Whether language models can be used to judge whether a document is relevant to a query — and what changes when they replace human assessors in that role. For decades, relevance judgments — the labels that say "this document is relevant to this query" — were produced by human assessors, expensive and slow but the gold standard for IR evaluation. Potthast's "LLMs for Relevance" line of work, run jointly with the Webis network, asked the practical question: how close can a language model get to expert human relevance judgments, and where does the substitution silently break? The findings are mixed in a productive way — close enough for some task types to enable evaluation at corpus scales human assessment cannot reach, and far enough off in others that any uncritical use produces systematically biased benchmarks. **Worth following when:** you want to know whether using an LLM to grade relevance is appropriate for your evaluation, and where it isn't. **Topics:** LLMs as substitutes for human relevance assessors; large-scale IR evaluation under labelling constraints; European open web-search infrastructure (OpenWebSearch.eu). **Key works:** Webis "LLMs for Relevance" line (2023 onward); long-standing PAN and TIRA contributions (with Stein); OpenWebSearch.eu open-infrastructure publications. --- ## Mengnan Du Source: https://ansmeter.com/whom-to-read/mengnan-du **Works on:** What kinds of explanations can be obtained for language-model outputs — and which of those explanations turn out to be reliable. The literature on explaining individual LLM outputs grew faster than the methods to validate those explanations against ground truth. Du's "Explainability for Large Language Models: A Survey" (2024, senior author with collaborators) sorts that landscape into categories that have come to matter — attention-based, gradient-based, perturbation-based, and natural-language explanations the model generates about itself, each with documented strengths and well-documented failure modes. The survey is particularly useful in the part most other surveys skip: case-by-case discussion of when an explanation method actively misleads rather than merely under-informs. **Worth following when:** you need to choose an explainability method for an LLM and want to know which of the options have been independently validated. **Topics:** taxonomy of LLM explainability methods; self-explanations by LLMs and their reliability; the methodology gap between proposing and validating explanations. **Key works:** "Explainability for Large Language Models: A Survey" (2024, senior author); ongoing TMLR Lab publications on trustworthy and explainable LLMs. --- ## Michihiro Yasunaga Source: https://ansmeter.com/whom-to-read/michihiro-yasunaga **Works on:** Whether a language model can correctly translate a natural-language question into a precise structured query — and what that translation reveals about reasoning over knowledge. Spider, which Yasunaga co-authored as part of a Yale-Stanford PhD collaboration, has been the canonical benchmark for text-to-SQL translation since 2018 — a task that looks deceptively simple ("turn this question into a SQL query") but actually requires the model to bind natural-language entities to schema columns, resolve aggregations, and structure joins correctly. The reason this matters for evaluating LLMs as reasoning systems is that text-to-SQL has a hard ground truth: the query either runs and returns the right rows, or it doesn't. His later work on QA-GNN and knowledge-graph-augmented reasoning extends the same posture — reasoning over structured knowledge that yields checkable answers. **Worth following when:** you want to evaluate LLM reasoning on tasks with strict, executable ground truth instead of human-judgment correctness. **Topics:** text-to-SQL evaluation (Spider); knowledge-graph-augmented question answering; reasoning tasks with verifiable correctness criteria. **Key works:** Spider text-to-SQL benchmark (2018, co-author); QA-GNN (2021, lead author); ongoing publications on reasoning over structured knowledge. --- ## Minlie Huang Source: https://ansmeter.com/whom-to-read/minlie-huang **Works on:** Whether the categories of "harmful" used to evaluate language-model safety transfer from Western to Chinese-language deployment contexts, where the regulatory frame and cultural categories are different. Most public LLM safety evaluation uses category schemes developed by Western labs — bias, toxicity, hate speech, jailbreak resistance, all benchmarked against English-language datasets and US/EU regulatory expectations. Huang's CoAI group at Tsinghua, with the more recent AISafetyLab framework, runs an explicit parallel for Chinese-language LLMs: what counts as "harmful" or "biased" output looks structurally different when the regulatory environment is different, the cultural taboos are different, and the deployed-model surface reaches users with different expectations. For Ansmeter, which evaluates engines that serve five language regions, this kind of locale-aware safety evaluation is the only honest version — pretending one safety taxonomy fits all five locales would be its own form of bias. **Worth following when:** you need to evaluate LLM safety in non-Anglophone deployment contexts, especially Chinese-language environments where the regulatory and cultural categories diverge. **Topics:** locale-aware LLM safety evaluation; dialogue-system safety and abuse detection; the Chinese-language open-LLM ecosystem (ChatGLM lineage). **Key works:** ChatGLM open language model series (2022 onward, key contributor); AISafetyLab safety-evaluation framework (2023, ongoing); CoAI group publications on conversational AI safety. --- ## Mohit Bansal Source: https://ansmeter.com/whom-to-read/mohit-bansal **Works on:** Whether the methods we use to evaluate language-only models still work when the same model has to handle images, speech, or other modalities at the same time — and what fails first when modalities are combined. Bansal's MURGe-Lab at UNC has produced one of the more sustained research lines on multimodal evaluation — vision-language reasoning, summarization faithfulness, long-form generation assessment — at a pace where most of the lab's papers identify a specific failure mode in existing evaluation methodology and propose a replacement. The body of work matters more than any single paper: it has shaped what subsequent generations of multimodal-evaluation papers consider rigorous. For Ansmeter, the question of how to evaluate brand-mention behavior in models that also process images (charts, product photos, screenshots) is precisely the multimodal-evaluation question Bansal's group has been carving out for several years. **Worth following when:** you need to evaluate a multimodal language model and want methodology informed by the longer arc of multimodal-evaluation research rather than ad-hoc extensions of text-only metrics. **Topics:** multimodal evaluation methodology (vision-language and beyond); summarization faithfulness assessment; the methodological practice of identifying failure modes before proposing metrics. **Key works:** MURGe-Lab body of publications on multimodal evaluation and reasoning (UNC, 2018 onward); summarization-faithfulness research line; ENGAGE NSF-AI Institute publications on multimodal AI. --- ## Nan Duan Source: https://ansmeter.com/whom-to-read/nan-duan **Works on:** Rewriting the user's query before retrieval so that the retriever has a chance of returning useful documents. A lot of real-world queries are bad queries — vague, missing context, phrased in ways the retriever cannot match against documents that would actually answer them. Duan's "Query Rewriting in RAG" line (2023, senior author) trains an LLM to rewrite the input query into a form the retriever can work with, and to do so adaptively based on what the system already has in context. The methodological consequence for Ansmeter is direct: when you measure a model's brand visibility, the upstream rewrite step is silently deciding what brands the retriever even sees as candidates — and that step is rarely audited. **Worth following when:** you want to know what happens between "user types a query" and "retriever looks for documents" — the layer where many brand-mention decisions are already made. **Topics:** LLM-based query rewriting for RAG; conversational query reformulation; the silent layer between user input and retrieval. **Key works:** "Query Rewriting in Retrieval-Augmented Large Language Models" (2023, senior author); foundational QA and multimodal-NLP work through the 2010s; ongoing foundation-model work in industrial research. --- ## Nathan Lambert Source: https://ansmeter.com/whom-to-read/nathan-lambert **Works on:** What happens to a language model between "trained on the internet" and "answering your question the way it does" — and how to study that step in public. Most of the post-training stack — RLHF, preference learning, instruction tuning — is developed inside closed labs and rarely described publicly. Lambert maintains an open counter-line: Tülu is the open post-training pipeline with published recipes; the RLHF Book is the only systematic treatment of the subject available without an NDA; his Interconnects newsletter is read across the industry as the primary open source on what actually goes on inside model labs during post-training. For Ansmeter this matters because post-training determines what the model says: pre-training fixes the corpus of knowledge, post-training decides which parts of it surface in any given answer. **Worth following when:** you want to understand why the same base model behaves very differently depending on which company shipped it — and where that behavior is decided. **Topics:** RLHF and preference learning in open documentation; post-training as the layer where model behavior is actually determined; open-model alignment pipelines (Tülu). **Key works:** Tülu post-training framework (2023, lead); RLHF Book (2024, ongoing); Interconnects newsletter as longer-form public analysis (2022, ongoing). --- ## Noah A. Smith Source: https://ansmeter.com/whom-to-read/noah-a-smith **Works on:** Whether a language model can produce its own training data — and what the methodological consequences are when that becomes the standard practice. Self-Instruct (2022, senior author) showed that an LLM can generate the instruction-tuning data needed to improve itself: prompt the model with a few seed examples of (instruction, input, output) tuples, ask it to generate more in the same pattern, filter for quality, and use the result as training data. The technique was immediately adopted across the industry — most current instruction-tuned LLMs were trained on data of this kind, much of it generated by GPT-3 or GPT-4. The methodological consequence for evaluation is uncomfortable: when models trained on synthetic data from older models are evaluated by judges that are also language models, the entire evaluation loop is increasingly endogenous to the same family of systems being measured. **Worth following when:** you want to think carefully about evaluation methodology in an era when training data, judges, and the systems being evaluated all come from related model lineages. **Topics:** self-generated instruction-tuning data (Self-Instruct); the methodological consequences of LLM-generated training data; the long arc of structured-prediction NLP. **Key works:** Self-Instruct (2022, senior author); long body of structured-prediction and probabilistic NLP work; ongoing UW and AI2 publications on responsible LLM development. --- ## Norbert Fuhr Source: https://ansmeter.com/whom-to-read/norbert-fuhr **Works on:** What information retrieval evaluation has been getting wrong, in writing, for the last several decades — and why each new generation of researchers makes the same mistakes. Fuhr's 2017 paper "Some Common Mistakes In IR Evaluation, And How They Can Be Avoided" reads like a checklist that every LLM-evaluation paper of the last three years should have been reviewed against. The list includes things like reporting averaged-over-runs metrics with no significance testing, comparing systems on different test collections, applying statistical tests inappropriate to the data — failures the IR field had already named and that the language-model-evaluation literature reinvented from scratch. His longer body of work, going back to probabilistic IR in the late 1980s, is the long story of trying to make retrieval a discipline rather than an engineering folklore. **Worth following when:** you suspect a current LLM-evaluation paper is making mistakes the IR community already documented two decades ago — and want a citation that backs you up. **Topics:** foundations of probabilistic IR; methodological mistakes in IR and NLP evaluation; the long-arc history of information retrieval as a discipline. **Key works:** "Some Common Mistakes In IR Evaluation, And How They Can Be Avoided" (2017); foundational work on probabilistic IR (1989 onward); Gerard Salton Award lecture and writings. --- ## Omer Levy Source: https://ansmeter.com/whom-to-read/omer-levy **Works on:** Whether a language model can infer the task from examples alone — without being told what to do — and what that ability reveals about how it represents instructions. Levy's "Instruction Induction" paper (2022, senior author) reversed the standard instruction-following experiment: instead of giving the model an instruction and observing whether it complies, give it a few input-output examples and ask the model to articulate what instruction those examples imply. The findings were both encouraging and disquieting — modern LLMs can induce reasonable instructions from a handful of examples on many tasks, which suggests their internal representation of "task" is genuinely abstract, but they sometimes induce instructions that subtly differ from the human-intended one in ways that affect downstream behavior. For Ansmeter, this matters because most engine queries don't include an explicit instruction — the engine is implicitly inferring the task from context, and the inference can go differently than expected. **Worth following when:** you want to understand the implicit-instruction-inference step that happens before an LLM produces an answer, especially when the user query is underspecified. **Topics:** instruction induction from examples; what implicit-instruction-inference reveals about model representations; the LIMA / Less Is More for Alignment line of work. **Key works:** Instruction Induction (2022, senior author); LIMA: Less Is More for Alignment (2023, co-author); foundational work on neural word embeddings (with Goldberg, 2014). --- ## Pang Wei Koh Source: https://ansmeter.com/whom-to-read/pang-wei-koh **Works on:** What happens when a language model encounters the kind of data it didn't see during training — and how to measure that gap rigorously, not just notice it after deployment. WILDS, which Koh co-led in 2021, was the first systematic benchmark suite for distribution shifts in machine learning — geographic shift, temporal shift, demographic shift, the kinds of real-world differences that quietly break models trained on convenient datasets. The result that organized the work was deflating: models with state-of-the-art performance on standard test sets routinely lost half their accuracy or more on the WILDS shifts, and no single robustness technique closed the gap. For Ansmeter this is the structural problem behind any benchmark of AI engines — what we measure on our query set is not what end-users encounter, and the gap deserves the same methodological care as accuracy itself. **Worth following when:** you suspect a benchmark result will not survive contact with real-world deployment data, and you want the literature that turned that suspicion into measurable shift categories. **Topics:** distribution-shift benchmarks (WILDS); foundation-model robustness under realistic conditions; the gap between in-distribution evaluation and deployment behavior. **Key works:** WILDS distribution-shift benchmark suite (2021, co-lead); foundation-model robustness publications (2021 onward); recent work on retrieval-augmented LMs and dataset attribution. --- ## Paolo Rosso Source: https://ansmeter.com/whom-to-read/paolo-rosso **Works on:** Building evaluation campaigns that work for Iberian-Romance languages — Spanish, Catalan, Portuguese — instead of porting English-centric methodology and accepting the resulting blind spots. Rosso has spent two decades organizing evaluation campaigns from inside the Spanish-language NLP community: IberLEF for the Iberian languages, PAN for plagiarism and authorship problems across languages, and a long line of shared tasks on author profiling, irony detection, hate-speech analysis, and disinformation. The substantive value is that the tasks were designed with native-speaker sensitivity to what linguistic and cultural features actually carry signal — irony in Spanish is not the irony of English-language datasets, and hate-speech categories shift with the regulatory and cultural context. For Ansmeter, which evaluates an engine's behavior in Spanish as one of its five language audits, this is the closest precedent for locale-native evaluation done with multi-year discipline. **Worth following when:** you need to evaluate language-model behavior in Iberian-Romance languages with methodology designed for those languages, not adapted from English. **Topics:** locale-native evaluation campaigns (IberLEF, PAN); irony, hate-speech, and disinformation detection in Spanish-language settings; multi-year shared-task organization as a research practice. **Key works:** IberLEF evaluation campaigns (2018 onward, co-organizer); long-standing PAN co-organization with Stein and Potthast (2009 onward); UPV PRHLT publications on multi-language NLP evaluation. --- ## Pascale Fung Source: https://ansmeter.com/whom-to-read/pascale-fung **Works on:** Whether the same language model performs the same kind of work across different languages — and whether evaluation methodology that's built for English is misleading us about what models do in everything else. Fung led the 2023 paper "Multitask, Multilingual, Multimodal Evaluation of ChatGPT" — one of the first systematic empirical studies of an LLM run through reasoning, hallucination, and interactivity tasks across multiple languages simultaneously. The findings were unflattering for the field: English-language evaluation results consistently overstated the model's competence in lower-resource languages, where the same prompts produced markedly worse reasoning quality and higher hallucination rates. For Ansmeter, which audits AI engines across five language regions, this is the cleanest published demonstration of why language-by-language evaluation is methodologically required, not an optional extension of an English baseline. **Worth following when:** you need to evaluate language-model behavior across multiple language environments and want the empirical literature that shows why per-language testing is non-negotiable. **Topics:** multilingual evaluation of LLMs (reasoning, hallucination, dialogue); empathetic and emotionally-aware dialogue systems; the methodological consequences of English-centric benchmark design. **Key works:** "Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity" (2023, senior author); CAiRE conversational AI publications; long-arc work on cross-lingual NLP since the 1990s. --- ## Percy Liang Source: https://ansmeter.com/whom-to-read/percy-liang **Works on:** How to measure language models so a measurement made today still means something next year. HELM came out of his group at Stanford — the closest thing the academic side has to a public, contestable reference scoreboard for language models. The choice that defines the work is that scoring functions, datasets, and prompt templates are all in the open, so you can disagree with the methodology in detail. That posture — methodology as the artifact, not scaffolding around a number — is what Ansmeter's own scoring tries to inherit. **Worth following when:** you want to read evaluation done by someone who treats the how of scoring as the contribution itself. **Topics:** holistic LLM evaluation; behavioral testing of foundation models; what benchmarks incentivize and what they hide. **Key works:** HELM (2022, ongoing); SQuAD reading-comprehension benchmark (2016, co-author); Foundation Model Transparency Index (2023, co-author). --- ## Peter Henderson Source: https://ansmeter.com/whom-to-read/peter-henderson **Works on:** Where the technical findings about language-model behavior actually matter — in audits, regulation, and legal liability — and what the gap between "we measured this" and "this changes what's allowed" looks like. Henderson trained both as a computer scientist and as a lawyer, which gives his published work an unusual posture: every empirical result about LLM behavior gets traced through to the regulatory or liability consequence it implies, or doesn't. His 2018 paper "Deep Reinforcement Learning that Matters" was an early demonstration that reproducibility failures in ML are not just an academic embarrassment but a basis for distrusting downstream claims that policy then has to act on. The POLARIS Lab at Princeton continues that line for the LLM era — foundation-model risk assessments, legal-compliance audits, and the kind of methodologically careful empirical work that actually survives use by regulators. **Worth following when:** you want technical AI-evaluation findings linked rigorously to the legal and policy decisions they should — and shouldn't — support. **Topics:** reproducibility failures and their regulatory consequences; foundation-model risk assessment from a law-and-CS dual perspective; AI policy grounded in empirical evaluation. **Key works:** "Deep Reinforcement Learning that Matters" (2018, lead author); ongoing POLARIS Lab publications on foundation-model legal compliance; legal-CS hybrid work on AI liability and regulation. --- ## Philip S. Yu Source: https://ansmeter.com/whom-to-read/philip-s-yu **Works on:** Bridging four decades of data-mining methodology to the question of how to evaluate large language models without reinventing techniques the field already has. Yu's published record covers most of what counts as data mining since the field had that name — stream mining, sequential pattern discovery, knowledge graphs, anomaly detection — produced from IBM Research in its long tenure and then continued at UIC, with several hundred patents along the way. His co-authorship of "A Survey on Evaluation of LLMs" reads as that body of work meeting the LLM-evaluation literature from above: many of the statistical and methodological questions that current LLM evaluation faces have been answered, partially, by techniques developed decades ago for similar problems on smaller-scale data. Reading him is the way to discover that some current "novel" evaluation methods are rediscoveries. **Worth following when:** you want LLM evaluation grounded in the longer history of data-mining methodology rather than treated as a discipline beginning in 2020. **Topics:** the long arc of data-mining methodology relevant to LLM evaluation; stream and sequential pattern mining; bridging legacy techniques to modern model evaluation. **Key works:** foundational work on stream mining and sequential pattern discovery (1990s–2000s); co-authorship of "A Survey on Evaluation of LLMs" (2024); decades of data-mining methodology publications at IBM Research and UIC. --- ## Prateek Mittal Source: https://ansmeter.com/whom-to-read/prateek-mittal **Works on:** How visual inputs become a new attack surface for safety-aligned language models that accept multimodal queries. The 2023 "Visual Adversarial Examples Jailbreak Aligned LLMs" paper from Mittal's Princeton group showed something the text-side safety community had not been paying attention to: a multimodal LLM that refuses a harmful text request will often comply with the same request when accompanied by an image carefully optimized to look unremarkable to humans but to push the model's internal representation into a region where its alignment training does not apply. The methodological consequence is concrete — any LLM safety evaluation that tests only text inputs is auditing a fraction of the actual attack surface. For Ansmeter the same lesson scales: brand-mention evaluation tested only in text mode misses the brand-mention behavior visible in the image-input or document-upload paths the engines also support. **Worth following when:** you want to evaluate LLM safety, robustness, or behavior across all the input modalities a modern engine actually accepts, not just the text-only slice. **Topics:** multimodal adversarial attacks on LLMs; safety alignment across input modalities; adversarial ML applied to deployed LLM products. **Key works:** "Visual Adversarial Examples Jailbreak Aligned LLMs" (2023, senior author); broader adversarial-ML and ML-security publications from Princeton. --- ## Qiang Yang Source: https://ansmeter.com/whom-to-read/qiang-yang **Works on:** Whether useful machine learning can happen when the data you'd train or evaluate on can't be moved to a single place — for legal, privacy, or commercial reasons. Yang co-authored the foundational textbook on federated learning, and his ongoing research line — first at HKUST and now also through WeBank — is about the same problem at industrial scale: build ML systems where multiple organizations contribute data without the data ever leaving their premises. The transfer-learning textbook before that set out the related question of when knowledge from one task or domain can usefully move to another. For Ansmeter, which currently audits engines on public-query behavior but may one day need to evaluate model behavior on data brand owners can't publish, federated-evaluation methodology is the literature that already worked out the privacy-versus-utility trade-offs. **Worth following when:** you need to evaluate or train ML systems on data that has to stay distributed, and want the methodology that the federated-learning community spent a decade developing. **Topics:** federated learning foundations; transfer learning across tasks and domains; the privacy-versus-utility trade-offs in distributed ML evaluation. **Key works:** "Federated Machine Learning: Concept and Applications" foundational paper (2019); transfer learning textbook (multiple editions); long body of work at HKUST and WeBank on federated AI. --- ## Quoc V. Le Source: https://ansmeter.com/whom-to-read/quoc-v-le **Works on:** The architectural and training-recipe building blocks that the modern LLM era was built on top of. Le's name appears as senior or co-author on a remarkable number of papers that turned out to be inflection points: Sequence-to-Sequence Learning with Neural Networks (2014), Doc2Vec, foundational AutoML and neural-architecture-search work, and — directly relevant to current LLM evaluation — Chain-of-Thought Prompting (2022, with Jason Wei and others). The throughline isn't a single research agenda but a habit of being early to the next architectural idea, often by a year or more. For Ansmeter, the practical implication is that almost every evaluation we run is downstream of one of his papers — Seq2Seq for the architecture, CoT for the prompting style most reasoning-aware evaluations now require. **Worth following when:** you want to trace a current LLM capability or evaluation practice back to its architectural origin and the trade-offs that shaped its design. **Topics:** sequence-to-sequence learning as the foundation of modern LLMs; chain-of-thought prompting; neural architecture search and AutoML. **Key works:** Sequence-to-Sequence Learning with Neural Networks (2014, co-author); Chain-of-Thought Prompting Elicits Reasoning in LLMs (2022, senior author); foundational AutoML and Neural Architecture Search publications. --- ## Rishi Bommasani Source: https://ansmeter.com/whom-to-read/rishi-bommasani **Works on:** What language-model developers are and aren't telling us about their own models — and how to measure that systematically. Most public LLM evaluation measures what a model can do. Bommasani's Foundation Model Transparency Index, run out of Stanford CRFM, runs along a perpendicular axis: how much we know about the model itself — what was in its training data, which internal evaluations were performed, which downstream impacts have been documented, what has been disclosed publicly versus kept private. The underlying argument is direct: capability benchmarks mean little when the conditions that produced those capabilities are not on the record, and treating transparency as a separate research discipline (rather than a PR concern) is the only way the comparison gets honest. **Worth following when:** you need to explain to an investor or partner why "this model has good benchmark scores" and "this model is transparent" are not the same answer. **Topics:** Foundation Model Transparency Index methodology; what model documentation should contain; the gap between what model labs say and what they make verifiable. **Key works:** Foundation Model Transparency Index (2023, ongoing, lead); "On the Opportunities and Risks of Foundation Models" report (2021, co-author); Stanford CRFM policy publications. --- ## Roi Reichart Source: https://ansmeter.com/whom-to-read/roi-reichart **Works on:** What "domain" means for a language model — when its training distribution stops matching its deployment context — and how to evaluate that mismatch rigorously. Reichart's longer research line has been domain adaptation in NLP: when a model trained on news articles is deployed on legal text, what specifically breaks, and how do you measure the breakage in a way that doesn't depend on the model itself being legible. As Co-Editor-in-Chief of the Transactions of the Association for Computational Linguistics, he has institutional influence on what the field accepts as rigorous evaluation methodology — TACL is one of the venues where evaluation papers either survive review or don't, and the standards there shape what subsequent work has to clear. For Ansmeter, which audits AI engines on queries the engines may not have been optimized for, the domain-adaptation literature is precisely the methodological backing. **Worth following when:** you want to evaluate a language model's behavior on inputs that may differ from its training distribution, and want the methodology that takes that mismatch seriously. **Topics:** domain adaptation methodology for language models; active learning under distribution shift; the editorial standards of TACL as a venue. **Key works:** body of work on NLP domain adaptation (2010s onward, Technion); long-term Co-Editor-in-Chief role at Transactions of the ACL; active-learning publications applied to language tasks. --- ## Seungone Kim Source: https://ansmeter.com/whom-to-read/seungone-kim **Works on:** Whether the field can have an open-weight LLM-evaluator alternative, so that "one black box judging another black box" stops being the only available option. The standard practice for scoring language models is to ask GPT-4 or another large closed model to grade the outputs — and the method works, while being methodologically shaky in a way that gets worse over time. Reproducing a result two years from now will be impossible, because the judge is a closed service with an opaque release cycle and regular weight updates. Kim's Prometheus and Prometheus 2 (lead author, 2023–24) are an attempt to give the community an open analogue: an evaluator with known weights and a documented training procedure, against which independent verification is actually possible. **Worth following when:** you plan to evaluate LLMs in research that has to remain reproducible, and you understand that using GPT-as-judge binds your science to someone else's product roadmap. **Topics:** open LLM-evaluators as alternatives to closed judges; what one transformer is actually comparing when it "scores" another; the limits of LLM-as-judge as a methodology. **Key works:** Prometheus (2023, lead author); Prometheus 2 (2024, lead author); CoT Collection (2023, lead author). --- ## Shafiq Joty Source: https://ansmeter.com/whom-to-read/shafiq-joty **Works on:** Evaluation methodology for language models when the deployment context is enterprise software rather than a research demo. Joty's earlier published work is in discourse-structure parsing — the kind of NLP where you ask whether a paragraph hangs together as an argument, not just whether each sentence parses. That habit carries into his more recent LLM-evaluation publications: an enterprise LLM application has to produce outputs that hold their argumentative shape across multi-turn use, and standard single-shot benchmarks don't capture that. The body of work reads as a steady reminder that "the model passes a benchmark" and "the model holds up under enterprise traffic" are not the same claim. **Worth following when:** you need an evaluation perspective informed by what shipping LLM features to enterprise customers actually demands of the underlying methodology. **Topics:** discourse structure as a lens on long-form LLM output; multi-turn evaluation beyond single-shot benchmarks; the gap between research benchmarks and enterprise deployment. **Key works:** earlier foundational work on discourse parsing (RST and beyond); ongoing publications on multi-turn LLM evaluation and reliability. --- ## Shayne Longpre Source: https://ansmeter.com/whom-to-read/shayne-longpre **Works on:** What's actually inside the data language models train on — and who can or can't tell. Longpre coordinates the Data Provenance Initiative, a collective audit of the licensing, sourcing, and legal status of the datasets behind modern LLMs. The project answers a question that turns out to be much harder than it sounds: when a model produces a sentence, what is the chain of evidence that the building blocks of that sentence were even in its training set? His earlier FLAN dataset work helped seed the instruction-tuning era; the current focus is closer to Ansmeter's transparency line — knowing what's upstream of an answer, not just what the answer looks like. **Worth following when:** you need to make defensible claims about what a model "knows" or "was trained on" for legal, audit, or product-positioning reasons. **Topics:** dataset licensing and provenance; the data lifecycle of generative models; what a real transparency report on training corpora should contain. **Key works:** Data Provenance Initiative (2023, ongoing, lead); FLAN dataset contributions (2021); audit reports on training-data licensing changes (2024, ongoing). --- ## Shinji Watanabe Source: https://ansmeter.com/whom-to-read/shinji-watanabe **Works on:** The open-source speech-processing infrastructure that lets academic and industrial groups train, evaluate, and compare voice-input or voice-output language systems on the same footing. Watanabe is one of the founders and primary maintainers of ESPnet (the End-to-End Speech Processing Toolkit), the open-source platform that became the default substrate for academic speech recognition and synthesis research over the past several years. The toolkit's design encodes a methodological position: evaluation should be reproducible across labs, baselines should run out of the box, and the tooling should make it harder to publish a "new SOTA" without comparing fairly to prior work. For Ansmeter, as the AI engines we evaluate add voice-mode interfaces, ESPnet is the methodological substrate that defines what fair speech-vs-speech comparison even looks like. **Worth following when:** you need to evaluate speech-input or speech-output capabilities of AI engines and want the open-toolkit methodology that the academic community converged on. **Topics:** open-source end-to-end speech processing (ESPnet); reproducibility-as-design in speech evaluation tooling; the methodology lineage for speech-recognition and speech-synthesis benchmarks. **Key works:** ESPnet end-to-end speech toolkit (2018 onward, co-creator and maintainer); body of work on end-to-end speech recognition and synthesis; CMU LTI publications on speech processing. --- ## Shuming Shi Source: https://ansmeter.com/whom-to-read/shuming-shi **Works on:** What "language model hallucination" looks like when you're responsible for shipping LLM-powered products to a billion-user surface. Shi co-authored the 2023 "Siren's Song" hallucination survey with Yue Zhang, but came at the problem from the deployment side — the NLP center of a major Chinese tech platform whose products reach roughly a billion users. The same taxonomy of hallucination types (factuality vs. faithfulness, intrinsic vs. extrinsic) reads differently when the question is operational: which categories actually show up in production traffic, which can be screened by a downstream filter, and which require changing the underlying training run. Worth reading alongside the academic side of the same survey because the deployment perspective surfaces failure modes that pure benchmark work misses. **Worth following when:** you want to know which hallucination categories break things in production versus which only matter in academic benchmarks. **Topics:** hallucination categories in deployed LLM products; industrial NLP at scale; the gap between benchmark behavior and production behavior. **Key works:** "Siren's Song in the AI Ocean: A Survey on Hallucination in LLMs" (2023, co-author); ongoing Tencent AI Lab NLP Center publications. --- ## Steven Schockaert Source: https://ansmeter.com/whom-to-read/steven-schockaert **Works on:** How to evaluate a retrieval-augmented system end-to-end when each part of it can fail in different ways. A RAG system has at least three places where it can go wrong: the retriever fetches the wrong context, the generator ignores the context it was given, or the answer is plausible-sounding but unsupported by what was retrieved. Schockaert's RAGAs framework (2023, senior author) gave the field its first set of automated metrics that disentangle these failure modes — faithfulness, answer relevance, context precision and recall — rather than collapsing them into a single quality score. The methodological point is closer to Ansmeter's own scoring philosophy than most evaluation frameworks: you don't get to call a system good if you can't say which part of it works. **Worth following when:** you need to evaluate a RAG pipeline and want to attribute failure to the retrieval, the generation, or the integration between them. **Topics:** automated RAG evaluation metrics; disentangling retrieval-side and generation-side failures; metric design for multi-stage pipelines. **Key works:** RAGAs: Automated Evaluation of Retrieval Augmented Generation (2023, senior author); ongoing Cardiff NLP work on RAG and knowledge-augmented LLMs. --- ## Tatsunori Hashimoto Source: https://ansmeter.com/whom-to-read/tatsunori-hashimoto **Works on:** Whether the numbers reported about language models are statistical findings or methodological artifacts. Hashimoto's group runs a quietly insistent line of inquiry: when a benchmark produces a result, what part of that result is the model and what part is the experimental setup? His 2023 benchmark for news summarization showed that LLMs declared "near human" on standard datasets owed as much to the brittle scoring conventions of those datasets as to model quality. The same skepticism drives his earlier work on distributional robustness — a model that does well on a held-out test set may have learned the test set's tilt rather than the underlying task. **Worth following when:** you want to read an evaluation paper and know whether you should believe the headline number. **Topics:** statistical evaluation of language models; benchmark validity and contamination; distributional robustness of trained models. **Key works:** Benchmarking LLMs for News Summarization (2023); foundational work on distributional robustness and group DRO (2019, ongoing). --- ## Tom Mitchell Source: https://ansmeter.com/whom-to-read/tom-mitchell **Works on:** Whether a language model has an internal representation of whether it's telling the truth — separable from what it actually outputs. Mitchell wrote the textbook the field grew up on (1997), founded the first Machine Learning Department in 2006, and in 2023 returned to first principles with a result that complicates the standard story about LLM hallucination: in his paper with Amos Azaria, a simple classifier trained on the hidden activations of an LLM predicts whether the model is about to produce a true or false statement, with accuracy well above chance. The model, in some readable sense, "knows" — and its surface behavior doesn't reflect that knowledge. **Worth following when:** you want a senior-figure take on truthfulness in LLMs that isn't downstream of the loudest current narratives. **Topics:** internal representations of truthfulness in LLMs; mechanistic probes for hallucination; the broader history of machine learning as a discipline. **Key works:** Machine Learning textbook (1997); "The Internal State of an LLM Knows When It's Lying" (2023, with Azaria); founding of the CMU Machine Learning Department (2006). --- ## Torsten Hoefler Source: https://ansmeter.com/whom-to-read/torsten-hoefler **Works on:** Generalizing language-model reasoning beyond linear chains of thought — into branching, backtracking, and recombination of intermediate reasoning steps. Chain-of-Thought prompting (2022) showed that a language model produces better answers when allowed to write its intermediate reasoning out loud, and Tree-of-Thoughts (2023) extended that into a branching search. Graph of Thoughts, which Hoefler's ETH group introduced in 2024, takes the next step — letting the model treat reasoning as a directed graph where intermediate states can be merged, scored, and backtracked from. The framework comes with the kind of systems-engineering instincts you'd expect from someone whose other day-job is architecting ML on a national supercomputer: tracking the computational cost of each reasoning expansion, not just its accuracy gain. **Worth following when:** you want to think about LLM reasoning as a structured search problem where compute spent on reasoning is a budget you have to allocate, not a free resource. **Topics:** structured reasoning over LLM intermediate states (Chain → Tree → Graph of Thoughts); compute-aware design of reasoning prompts; ML systems engineering at HPC scale. **Key works:** Graph of Thoughts (2024, senior author); broader work on ML systems and scaling at ETH Zürich and CSCS; HPC-side contributions to large-scale model training infrastructure. --- ## Tushar Khot Source: https://ansmeter.com/whom-to-read/tushar-khot **Works on:** What "reasoning ability" actually means as something you can put on a benchmark — and what changes when you also try to coach the model to do reasoning through structured prompting. ARC, which Khot helped design at AI2, is the grade-school-science benchmark that current LLM papers cite as evidence of "reasoning ability" — questions where the model has to work out which scenario explains an observation, a step that goes beyond surface fact retrieval. The same instinct drives his Decomposed Prompting work (2023, lead author): if a complex task can be broken into a stable set of sub-tasks, you can prompt each sub-task in isolation and recompose the results, getting both better accuracy and an inspectable trace of what the model did. Together, the two lines make ARC a stronger evaluation tool — you can identify exactly where in the reasoning chain the model lost the question. **Worth following when:** you want benchmark design and prompting methodology treated as the same problem in LLM evaluation. **Topics:** reasoning benchmarks (ARC and successors); decomposed prompting and reusable sub-task patterns; reasoning-trace inspection as part of evaluation. **Key works:** AI2 Reasoning Challenge / ARC (2018, co-author); Decomposed Prompting (2023, lead author); ongoing AI2 publications on machine reasoning evaluation. --- ## Wayne Xin Zhao Source: https://ansmeter.com/whom-to-read/wayne-xin-zhao **Works on:** The synthesis side of LLM research — what it takes to read every paper of the moment and produce something other researchers can navigate. Zhao led the writing of "A Survey of Large Language Models" — the 100+ page synthesis paper that became, for a stretch of 2023, the most-cited reference for what the LLM field actually contained. The work behind that kind of paper is less visible than the methodological papers that produced the surveyed work: someone has to read everything, identify which claims actually survive scrutiny when laid next to each other, and structure the result so it remains useful when the field moves another six months. His longer research line in information retrieval and recommender systems gives the survey work its underlying methodological discipline — the same instincts that catch failure modes in recommender-eval transfer cleanly to catching them in LLM-eval claims. **Worth following when:** you want a synthesis-style read on the LLM landscape from someone with strong IR and recommender-systems foundations to identify which claims hold up under scrutiny. **Topics:** comprehensive LLM landscape synthesis; information retrieval and recommender systems foundations; the methodological discipline of survey writing. **Key works:** "A Survey of Large Language Models" (2023, lead first author); long body of work on IR and recommender systems; Renmin GSAI publications on LLM-augmented retrieval. --- ## Weijia Shi Source: https://ansmeter.com/whom-to-read/weijia-shi **Works on:** Making retrieval-augmented generation work when the language model itself is a closed box you can't fine-tune. The dominant assumption in retrieval-augmented systems is that you train the retriever and the generator together, or at least fine-tune one to fit the other. REPLUG (2023, lead author) showed that this isn't necessary — the retriever can be trained against a closed-source language model treated as a frozen black-box scoring function, and the resulting system improves the LLM's outputs without ever touching its weights. The methodological consequence is that retrieval-augmentation research stopped being something that only applied to open-weight models, which is the only condition under which it's relevant for evaluating systems like the major commercial AI engines. **Worth following when:** you want to understand how retrieval improves closed-source LLMs that don't expose their weights or training signals. **Topics:** black-box retrieval-augmented generation; retriever training without LM access; the architectural assumptions buried in earlier RAG research. **Key works:** REPLUG: Retrieval-Augmented Black-Box Language Models (2023, lead author); In-Context RALM (2023, co-author). --- ## Wen-tau Yih Source: https://ansmeter.com/whom-to-read/wen-tau-yih **Works on:** Making the retrieval step inside retrieval-augmented systems good enough that the generation step has something to work with. Long before RAG existed as a term, Yih was building the retrieval substrate it depends on. Dense Passage Retrieval (2020) is the paper that made dense-vector search competitive with traditional inverted-index methods for open-domain QA — virtually every modern RAG system descends from that result, including the ones that don't credit it. Reading Yih is the cleanest way to understand why the retrieval in retrieval-augmented generation is usually where systems silently break, and why FActScore (which he co-authored) bothers to score atomic facts against an external knowledge source rather than against the model's parametric memory. **Worth following when:** you want to understand why a RAG-augmented model still hallucinates — and whether the failure is in retrieval or in generation. **Topics:** dense passage retrieval; open-domain question answering as a discipline; what factual grounding does and doesn't fix. **Key works:** Dense Passage Retrieval / DPR (2020, co-author); WikiQA (2015, co-author); FActScore (2023, co-author). --- ## Wenjie Li Source: https://ansmeter.com/whom-to-read/wenjie-li **Works on:** How generative retrieval relates to the other generation tasks — summarization, question answering — that the same system architecture has to handle. Generative retrieval was first treated as a niche IR technique: train a sequence-to-sequence model to emit document identifiers given a query. Li's research, coming from a summarization and NLG background, reframed the picture — once a system can generate document IDs, it can also generate the summary of those documents conditioned on the query, and the answer to the query conditioned on the documents, all from the same architecture and weights. For Ansmeter, this matters because the boundary between "the system retrieved sources" and "the system summarized them" gets blurry: a modern engine in answer-with-citations mode is doing both, and Li's lineage of work is the cleanest place to see why the two tasks were always entangled. **Worth following when:** you want to think about retrieval, summarization, and question-answering as a single generation problem rather than as a pipeline of separate components. **Topics:** generative document retrieval; the unification of summarization and retrieval as generation tasks; query-conditioned summarization in RAG-adjacent systems. **Key works:** Neural Corpus Indexer (NCI, 2022, co-author) and downstream generative-retrieval line; long body of work on text summarization and query-focused NLG; HK PolyU Cognitive Computing Lab publications. --- ## Xia Hu Source: https://ansmeter.com/whom-to-read/xia-hu **Works on:** Whether the academic state of the art in language models can be turned into something a practitioner — not a research lab — can actually deploy and trust. Hu's "Harnessing the Power of Large Language Models in Practice" survey (2023, senior author) sits in a different corner of the survey literature from the academically dense Zhao-or-Wen-style overviews: it organizes the field around questions a real deployment team has to answer — which model fits which problem class, what the data requirements look like at production scale, what fine-tuning versus prompting actually costs, where the trust gaps appear. The same instinct runs through his D2K Lab work on automated and interpretable ML — methodology designed so that someone outside the research lab can reproduce the result and audit how it was reached. For Ansmeter, this is the literature that maps how brand-mention behavior should be evaluated by someone who has to defend the evaluation to a procurement team. **Worth following when:** you need to evaluate LLMs from a practitioner's stance and want the literature framed for deployment rather than for benchmark records. **Topics:** practical LLM deployment surveys; automated and interpretable ML; LLM efficiency considerations for production use. **Key works:** "Harnessing the Power of Large Language Models in Practice: A Survey" (2023, senior author); D2K Lab automated-ML publications; long body of work on interpretable ML and AutoML at Rice. --- ## Xing Xie Source: https://ansmeter.com/whom-to-read/xing-xie **Works on:** Connecting the recommender-systems tradition of measuring user-facing AI behavior with the new evaluation challenges modern LLMs pose. Xie's published career has crossed several adjacent fields — recommender systems, spatial data mining, responsible AI infrastructure — that all had to answer the same operational question: what does this AI system actually do for the user, and how do you measure that against the user's actual interests rather than against a benchmark you defined. His co-authorship of the 2023 "Survey on Evaluation of LLMs" reads as that long career meeting the LLM moment: many of the methodological frames for measuring fairness, drift, or unintended user-facing effects in recommender systems transfer directly to language models, often without modification. For Ansmeter, which evaluates how language models shape what users hear about brands, this is the closest precedent literature — recommender-system evaluation methodology applied to a new substrate. **Worth following when:** you want LLM evaluation methodology informed by the longer arc of evaluating user-facing AI systems before LLMs existed. **Topics:** recommender-systems evaluation methodology applied to LLMs; responsible-AI infrastructure inside large research labs; the bridge between recommender-system metrics and LLM evaluation. **Key works:** co-authorship of "A Survey on Evaluation of LLMs" (2023, with Wang and collaborators); long body of work on recommender systems and urban computing at Microsoft Research Asia; ongoing responsible-AI publications. --- ## Xipeng Qiu Source: https://ansmeter.com/whom-to-read/xipeng-qiu **Works on:** Building an open Chinese-language large language model that the academic community can actually study under the hood — and the tooling around it. Qiu's group at Fudan released MOSS, one of the early open Chinese-language LLMs that came with enough documentation and training tooling for independent groups to reproduce and modify. The same lab is responsible for FastNLP — one of the more widely used Chinese-side NLP frameworks for sequence modeling — and a research line on in-context demonstration retrieval that connects to current evaluation methodology. For Ansmeter, which audits AI engines including Chinese-language behavior, MOSS sits in the part of the literature that documents what an open Chinese LLM actually looks like inside, and Qiu's group is the most readable entry point. **Worth following when:** you need to understand the open Chinese-language LLM landscape from inside its academic origin, rather than from English-language news coverage of it. **Topics:** open Chinese-language LLMs (MOSS); the FudanNLP / FastNLP ecosystem; in-context demonstration retrieval methodology. **Key works:** MOSS open Chinese LLM (2023, project lead); FastNLP / FudanNLP framework body of work; publications on in-context learning and demonstration retrieval. --- ## Xueqi Cheng Source: https://ansmeter.com/whom-to-read/xueqi-cheng **Works on:** What information retrieval research from inside the Chinese IR tradition has been arguing — and how its angle on generative retrieval differs from the Western canon. Cheng has been one of the central figures in Chinese IR for over two decades, with the kind of multi-decade publication arc that crosses social-network analysis, web data, and retrieval, and which has lately converged on the generative-retrieval line his group at ICT-CAS publishes. The work differs in emphasis from the equivalent Western literature: more weight on graph and network structure as a retrieval signal, more attention to scaling under the specific data and user-behavior conditions of the Chinese-language web. For Ansmeter — which audits AI engines that operate in language environments shaped by very different web ecosystems — reading Cheng's line is one of the most useful correctives against treating "IR research" as English-web research wearing a translation hat. **Worth following when:** you want to know what retrieval research looks like outside the Anglophone canon, especially the part that focuses on graph and network structure and on non-English-web scale. **Topics:** generative retrieval from the ICT-CAS school; graph and network structure as retrieval signal; the long arc of Chinese IR as a research tradition. **Key works:** body of work on graph-based and network-aware IR (1990s onward, ICT-CAS Beijing); ongoing generative-retrieval publications from his group; SIGIR and CIKM best-paper output across two decades. --- ## Yang Liu Source: https://ansmeter.com/whom-to-read/yang-liu **Works on:** Whether a more capable language model can be used to grade the outputs of a less capable one — and how to do that without fooling yourself. The 2023 G-Eval paper crystallized a practice that had been creeping into the field for a year before: instead of writing yet another automatic metric for NLG quality, use a strong LLM as the judge, prompt it with a structured rubric, and read its output as a score. G-Eval correlated with human judgment better than every prior automatic metric on the same tasks, and the paper is now the citable reference for almost any system that scores generated text with a model — including Ansmeter's, which depends on this approach existing as a defensible methodology rather than a clever hack. **Worth following when:** you want to understand the methodological foundation under "let the model grade it" and how to argue for it in a peer-reviewed setting. **Topics:** LLM-as-judge as a scoring methodology; chain-of-thought prompting in evaluation; human-alignment of automatic NLG metrics. **Key works:** G-Eval: NLG Evaluation Using GPT-4 with Better Human Alignment (2023, lead author); subsequent LLM-as-judge methodology refinements. --- ## Yann LeCun Source: https://ansmeter.com/whom-to-read/yann-lecun **Works on:** What machine-learning systems should look like, set against what they currently are. LeCun shares the 2018 Turing Award for the work that made deep learning workable at scale. The position he has held publicly for the past several years — that autoregressive next-token prediction is a dead end as a path toward anything resembling general reasoning — is one of the few senior dissents from current LLM consensus that is argued in technical architectural detail. Reading him is useful even when you suspect he's wrong, because the disagreement is specific: world models, joint-embedding predictive architectures, energy-based learning, all set against why he thinks the current generation of LLMs cannot get there from here. **Worth following when:** you want a senior voice arguing against current LLM consensus in concrete architectural terms rather than in vague philosophical objection. **Topics:** self-supervised learning; the limits of autoregressive generation; world-model and joint-embedding approaches to learning systems. **Key works:** Turing Award–era foundational deep-learning contributions (shared 2018); ongoing JEPA / world-model line (2022 onward); public-facing critiques of the autoregressive LLM paradigm (2022, ongoing). --- ## Yarin Gal Source: https://ansmeter.com/whom-to-read/yarin-gal **Works on:** Quantifying when language models don't know what they're saying. Gal worked out the foundations of Bayesian uncertainty in deep neural networks while that was still a small literature, and has spent the past few years pointing the same instrument at LLMs. His 2024 Nature paper on semantic-entropy hallucination detection is the throughline made concrete: a fluent answer and a true one are different things, and the gap is something you can measure from the outside rather than ask the model to self-report. The method does not require labeled ground truth or access to model internals, which is the only kind of method that actually applies to a closed commercial LLM. **Worth following when:** you need to take a language-model output and ask "should this be trusted?" — and you want an answer that doesn't depend on the model agreeing. **Topics:** semantic uncertainty in LLMs; hallucination detection without labeled ground truth; how AI safety institutes translate uncertainty research into oversight. **Key works:** "Dropout as a Bayesian approximation" (2016); "Detecting hallucinations in LLMs using semantic entropy" (Nature, 2024); ongoing Oxford OATML group output. --- ## Yejin Choi Source: https://ansmeter.com/whom-to-read/yejin-choi **Works on:** What it would take for a language model to have what humans have in spades and what LLMs reliably lack — common sense. Choi has been the most consistent voice in NLP arguing that scale alone does not produce common-sense reasoning, and her group has built the benchmarks to prove it: HellaSwag (where humans get 95% right and language models from 2019 got 47%), PIQA for physical commonsense, SIQA for social, COMET for commonsense knowledge graphs. The framing she brought to the field — that current evaluation gives models credit for surface-pattern matching on tasks where humans are using actual world models — has aged well as larger LLMs improved on the benchmarks while still producing reasoning failures the benchmarks were designed to expose. For Ansmeter, this is the literature that justifies our skepticism of headline benchmark numbers: a model scoring 92% on commonsense benchmarks may still be doing something different from what humans do to score 95%, and the difference matters for downstream behavior. **Worth following when:** you want LLM evaluation methodology grounded in the gap between benchmark performance and actual reasoning competence — from someone who has been making that argument since before it was fashionable. **Topics:** commonsense reasoning evaluation (HellaSwag, PIQA, SIQA, COMET); the methodological gap between benchmark scores and reasoning competence; long-arc work on what NLP evaluation actually measures. **Key works:** HellaSwag (2019, senior author); PIQA: Reasoning about Physical Commonsense (2020, co-author); COMET: Commonsense Transformers for Automatic Knowledge Graph Construction (2019, senior author); Defending Against Neural Fake News (Grover, 2019, senior author); AI2 Mosaic Team body of work on commonsense AI. --- ## Yoav Goldberg Source: https://ansmeter.com/whom-to-read/yoav-goldberg **Works on:** Whether the chain of reasoning a language model produces is the chain of reasoning it actually followed. Goldberg wrote one of the standard NLP textbooks for the neural era (2017) and has spent the years since pushing on a question that gets harder as models get more capable: is the explanation an LLM gives for its answer the actual cause of that answer, or a plausible-sounding cover story produced after the fact? His work on faithfulness in generated explanations is a steady reminder that "the model said why it did X" and "we know why the model did X" are different claims, and most current evaluation methodology conflates them. **Worth following when:** you want someone who treats LLM-generated explanations with the same skepticism applied to the answers themselves. **Topics:** faithfulness of generated explanations and chains of thought; the limits of post-hoc model interpretability; NLP as a field — its history and its current confusions. **Key works:** Neural Network Methods for NLP (2017); ongoing publications on faithfulness in NLG; public NLP commentary as a longer body of work. --- ## Yoav Shoham Source: https://ansmeter.com/whom-to-read/yoav-shoham **Works on:** What can be measured about the state of AI from outside any single company — and what retrieval-augmentation looks like when you don't need to retrain anything. Shoham's career spans multiple eras of AI — multi-agent systems and game theory in the textbook he co-authored, then a long stretch chairing the Stanford AI Index, an annual measurement of the field that has become the closest thing the industry has to an official accounting. The "In-Context Retrieval-Augmented Language Models" paper (2023, senior author) sits inside the AI21 stack but reads as a method paper: retrieval-augmentation works on frozen LLMs with off-the-shelf retrievers, no joint training needed, which is exactly the regime that applies when you can't open up the model you're trying to improve. The two lines — measurement of the field and methods that work in the closed-model regime — meet at a position Ansmeter shares: you don't get to claim industry-level insight without admitting what you can and can't measure. **Worth following when:** you want a senior AI figure who treats "measuring the state of the field" and "making methods that work without model access" as one continuous project. **Topics:** annual industry measurement (Stanford AI Index); in-context retrieval-augmentation without joint training; multi-agent systems as a foundational subfield. **Key works:** Multiagent Systems textbook (2009, with Leyton-Brown); In-Context Retrieval-Augmented Language Models (2023, senior author); Stanford AI Index annual report (chair). --- ## Yonatan Belinkov Source: https://ansmeter.com/whom-to-read/yonatan-belinkov **Works on:** What's actually inside a language model's hidden representations — and which of those internal states map onto things humans would recognize as knowledge. Long before mechanistic interpretability became a mainstream subfield, Belinkov was running probing experiments on neural language models — training small classifiers to predict whether the model's hidden states encode part-of-speech, syntactic structure, semantic role, factual knowledge. The early results were sobering and stabilizing: models often did encode linguistic structure in measurable ways, but the encoding was distributed and entangled in patterns that contradicted the assumption that a single layer or unit "represents" a single concept. The line continues into his current LLM-era work, where the question is whether the same probing techniques scale to models a thousand times larger. **Worth following when:** you want to ground claims about what a language model "knows" in something more measurable than its output behavior. **Topics:** probing classifiers for linguistic and factual knowledge in LMs; the structure of distributed internal representations; the BlackboxNLP tradition. **Key works:** foundational probing experiments on neural LMs (2017 onward); "Linguistic Knowledge and Transferability of Contextual Representations" (2019, co-author); co-founding of the BlackboxNLP workshop series. --- ## Yue Zhang Source: https://ansmeter.com/whom-to-read/yue-zhang **Works on:** Organizing what the field calls "hallucination" into categories that actually mean different things. Before 2023, "LLM hallucination" was a single bucket — a wrong fact, a made-up citation, a confused chain of reasoning, all got the same word. Zhang's "Siren's Song in the AI Ocean" survey separated the categories with the discipline of a taxonomist: factuality hallucinations (the model contradicts the world), faithfulness hallucinations (the model contradicts its own context), and several finer subcategories that have since become standard vocabulary in the literature. Reading him is the fastest way to stop saying "hallucination" as if it were one phenomenon. **Worth following when:** you need to talk about LLM hallucination with technical precision rather than as a single hand-wave term. **Topics:** taxonomy of LLM hallucination; faithfulness versus factuality; survey methodology in NLP. **Key works:** "Siren's Song in the AI Ocean: A Survey on Hallucination in LLMs" (2023); ongoing Westlake NLP group output on structured prediction and LLM evaluation. --- ## Yulia Tsvetkov Source: https://ansmeter.com/whom-to-read/yulia-tsvetkov **Works on:** What kinds of bias and harm look like in language-model outputs across languages — especially the languages and communities the field's standard evaluation has historically ignored. The standard pipeline for studying bias in language models has been English-first by default: train on English, evaluate on English, draw conclusions, then port to other languages as an afterthought. Tsvetkov's UW group runs the inverse pipeline — start with the languages and communities that English-centric methods don't reach, ask which bias and harm categories actually surface in those settings, and let the multilingual data shape the analytical categories from the start. For Ansmeter, which evaluates engines in five language regions, this is the methodological backing for treating each locale's bias profile as its own object of study. **Worth following when:** you need to evaluate model bias or harm in a non-English context and want the literature that takes multilingual ethics as a primary concern. **Topics:** multilingual bias and harm evaluation; low-resource language NLP; ethics in NLP research as it intersects with language diversity. **Key works:** "Demystifying Prompts in Language Models via Perplexity Estimation" (2023, co-author); long body of work on low-resource and multilingual NLP; UW publications on bias evaluation across languages. --- ## Zhaochun Ren Source: https://ansmeter.com/whom-to-read/zhaochun-ren **Works on:** Whether a language model can be trusted with the job that's currently done by an information retrieval system — and which parts of that job it actually does well. "Is ChatGPT Good at Search?" (2023, with Ren as senior author) is one of the cleanest empirical studies of whether large language models can replace traditional ranking components in a retrieval pipeline. The answer turns out to be: for re-ranking a small candidate set, yes, quite well; for first-stage retrieval over a large corpus, no, not really — and the gap between those two tasks is one most LLM-centric papers gloss over. Ren's broader work in conversational IR pushes the question further: when retrieval happens inside a multi-turn conversation, the system is doing several different things at once, and they should not all be benchmarked the same way. **Worth following when:** you want empirical results on where LLMs can replace traditional IR components and where they break down. **Topics:** LLMs as re-rankers in retrieval pipelines; conversational information retrieval; the difference between candidate-set re-ranking and corpus-scale first-stage retrieval. **Key works:** "Is ChatGPT Good at Search? Investigating LLMs as Re-Ranking Agents" (2023, co-author); ongoing Leiden LIACS publications on conversational IR and RAG. --- ## Zhiting Hu Source: https://ansmeter.com/whom-to-read/zhiting-hu **Works on:** Treating a language model as one component in a larger planning system — instead of asking it to do reasoning end-to-end in its own head. The dominant frame for LLM reasoning is to keep everything inside the language model: prompt it to produce reasoning steps, sample from the resulting distribution, hope the right answer falls out. RAP (Reasoning as Planning, 2023, with Hu as senior author) makes a different bet — use the LLM as a world model that predicts what happens after a candidate action, and combine it with classical Monte Carlo Tree Search to choose which action to actually take. The architectural consequence is that reasoning becomes auditable in a way pure prompting isn't: you can inspect the search tree, see which states the model evaluated as promising, and trace why the final answer was chosen. **Worth following when:** you want to know how LLM reasoning behavior changes when the model is constrained to operate inside an explicit search structure rather than freelancing through prompted text. **Topics:** language models as world models for planning; integration of classical search (MCTS) with LLM components; auditable reasoning traces in agentic systems. **Key works:** RAP — Reasoning with Language Model is Planning (2023, senior author); broader publications on combining symbolic and neural methods for structured generation. ---