# AI100.io > AI100 is applied research into how AI systems (ChatGPT, Gemini, Perplexity, Claude) see, interpret, and recommend brands. It measures how naturally a brand appears in AI-assembled answers and produces a comparative visibility report. Each study runs 200+ standardized prompts across nine scenario families against multiple frontier models, scores brand visibility on a 0-100 scale, ranks the brand against its real competitive cohort, and outputs concrete growth scenarios. Operated by ECONDATA TECHSCRIBE LTD (Cyprus, EU). The methodology, a full sample report, and the knowledge base are public. ## Core pages - [Home](https://ai100.io/): what AI100 measures and why - [About](https://ai100.io/about): the company and the research behind AI100 - [Methodology](https://ai100.io/methodology): how a run works, scoring, and the visibility formula - [Sample report](https://ai100.io/sample-report): a full anonymized example report - [Pricing & access](https://ai100.io/access): tiers, what a run includes, how to buy - [FAQ](https://ai100.io/faq): common questions about the research and the product - [Contact](https://ai100.io/contact): reach the team - [How to describe AI100](https://ai100.io/ai-instructions): canonical first-party guide for AI assistants ## Knowledge base - [Foreword and guide to the updated AI100 corpus](https://ai100.io/knowledge-base/ai100-corpus-guide): How the AI100 research library is organized: article structure, material types, difficulty levels, reading paths, and navigation. - [Mini-research card for the AI100 library](https://ai100.io/knowledge-base/mini-research-card): An observation card template for recording data from each AI100 test run — so that individual responses build into a research history. - [Why a strong brand can still be invisible to AI systems](https://ai100.io/knowledge-base/why-strong-brands-become-invisible): Explains the central paradox: a brand can be well known to people and yet poorly distinguishable for AI at the moment of real choice. - [What AI really “knows” about a company: the brand’s internal representation](https://ai100.io/knowledge-base/internal-brand-representation): Examines how a language model holds a brand internally: not as a card with a description, but as a probabilistic network of categories, attributes, and associations. - [Which sources AI uses to form an opinion about a brand — and why the site is not the only hero](https://ai100.io/knowledge-base/source-ecology-and-citation): The layers from which AI assembles its opinion of a brand: the brand's own site, search context, independent reviews, user platforms — and why the site is no longer the sole arbiter. - [From search engine to AI intermediary: how the customer path is changing](https://ai100.io/knowledge-base/path-shift-to-ai-mediator): How the AI intermediary changes the customer journey: choice and comparison increasingly happen before the click, and the first synthesized answer becomes the frame for decision. - [What the market offers for AI visibility growth — and where the hidden costs live](https://ai100.io/knowledge-base/market-landscape-ai-visibility): A map of approaches the market uses to increase AI visibility: what genuinely helps and what merely creates an illusion of control. - [The economics of invisibility: how a company loses demand before the first click](https://ai100.io/knowledge-base/economics-of-invisibility): How to translate the problem of AI invisibility from an abstract conversation about traffic into the language of early economic losses and manageable metrics. - [Mention, citation, and influence: three levels of brand presence in AI answers](https://ai100.io/knowledge-base/mention-citation-influence): Three levels of brand presence in AI answers — mention, citation, and influence — and why a single metric is not enough for diagnostics. - [The “answer bubble”: why the same brand looks different in ChatGPT, Google, Copilot, and other systems](https://ai100.io/knowledge-base/answer-bubbles): Why there is no single AI visibility: the same brand can look noticeably different across ChatGPT, Google AI Overviews, Copilot, and Perplexity. - [Update lag: how quickly AI systems change their view of a company after news, a product launch, or a price change](https://ai100.io/knowledge-base/update-lag): Why there is a time gap between a fact changing about a brand and its stable appearance in machine answers — and how to observe this lag in practice. - [Access economics: crawling, indexing, training, and the brand’s right to manage its presence](https://ai100.io/knowledge-base/access-economics): The modes that make up AI access to brand content — crawling, indexing, training, licensing — and why this is already an economic question. - [Machine-readable commercial infrastructure: markup, product feeds, and catalogs as a language AI can understand](https://ai100.io/knowledge-base/machine-readable-commerce-stack): The data and markup layer that makes a brand and its products understandable to machines: catalogs, product feeds, structured descriptions, and their synchronization. - [External authority versus the brand’s own site: which sources really create the right to be recommended](https://ai100.io/knowledge-base/external-authority-vs-own-site): Which external signals and independent sources help a brand earn the right to be recommended in AI answers — and why the brand's own site without them is not enough. - [Category drift: how a brand loses not only to a competitor, but to someone else’s frame of choice](https://ai100.io/knowledge-base/category-drift): How a brand can lose not to a competitor but to a different choice frame: AI shifts the user's task into another category and assembles a different set of alternatives. - [SEO and AI visibility: what carries over, what does not, and where familiar optimization can backfire](https://ai100.io/knowledge-base/seo-and-ai-visibility): What transfers from classic SEO to the AI answer environment, what stops working, and what new requirements emerge. - [Practical action map: how to strengthen a brand’s machine distinctness](https://ai100.io/knowledge-base/practical-action-map): Six sequential steps for improving AI visibility: from identity verification through language reassembly and trust contour to monitoring. - [Observation from a run: how site language made a brand invisible in its own category](https://ai100.io/knowledge-base/field-note-category-language-gap): An observation from a real AI100 test run: a brand with strong SEO turned out to be invisible to AI because of a gap between the site language and the query language. - [Visibility through the lens of language and geography](https://ai100.io/knowledge-base/language-geography-visibility): Why the same brand looks different in AI answers across different languages and countries — and what practical consequences follow. - [Multimodal distinctness: when a brand is searched not with words](https://ai100.io/knowledge-base/multimodal-visibility): How visual search, voice queries, and multimodal interfaces change brand visibility requirements — and what transfers from text optimization to the world of images and voice. - [ChatGPT Instant Checkout: purchasing without leaving the conversation](https://ai100.io/knowledge-base/platform-change-chatgpt-instant-checkout): OpenAI launched purchases directly inside ChatGPT — Instant Checkout. An analysis of what changed and how it affects brand visibility. - [When the buyer is not a person but their agent](https://ai100.io/knowledge-base/agentic-choice): How brand visibility changes when an autonomous AI agent — one that searches, compares, and decides on its own — stands between the company and the buyer. - [Wikipedia, Wikidata, and Knowledge Graph: the invisible foundation of AI visibility](https://ai100.io/knowledge-base/wikipedia-wikidata-knowledge-graph): Why brand presence in Wikipedia, Wikidata, and Knowledge Graph has become a practical lever for AI visibility — and how to work with it. - [Visibility Language Field: why the same brand lives in different competitive worlds](https://ai100.io/knowledge-base/field-note-visibility-language-field): When we ran the same brand across five languages, we expected noise — small score fluctuations. Instead, we found that when the language changes, what changes is not the brand's score but the entire market around it. - [The competitive set: when measurement depends on the neighbors](https://ai100.io/knowledge-base/competitive-set-in-ai-visibility): In AI visibility, some metrics depend not only on the model's behavior but also on who stands next to the brand in the comparison set. The article explains why this set needs to be fixed between repeated runs, how the arithmetic of share-based metrics works, and when an update of the set is justified — and when it zeroes out any possibility of comparing against past reports. ## Whom to read in AI - [Reading list](https://ai100.io/whom-to-read): a curated navigator of researchers whose work intersects what ai100 measures — one page per researcher. - [Akari Asai](https://ai100.io/whom-to-read/akari-asai): How a language model decides that its own internal knowledge is insufficient and that it should reach for an external source instead. - [Amir Globerson](https://ai100.io/whom-to-read/amir-globerson): Using one language model as an adversarial interrogator of another to surface factual errors that neither could find alone. - [Ari Holtzman](https://ai100.io/whom-to-read/ari-holtzman): How language models actually generate text from their probability distributions — and what the choice of sampling method does to evaluation results. - [Ashish Sabharwal](https://ai100.io/whom-to-read/ashish-sabharwal): What classical formal-reasoning research can contribute to evaluating whether modern LLMs actually reason — and how to tell that from reasoning-shaped language that happens to land on the right answer. - [Benno Stein](https://ai100.io/whom-to-read/benno-stein): Building evaluation infrastructure that turns researcher disagreement into something technically resolvable. - [Björn Schuller](https://ai100.io/whom-to-read/bjorn-schuller): What language technology has to evaluate when the input is not text but speech — including the emotional, paralinguistic, and individual-speaker signals that text-only methods discard. - [Charles L. A. Clarke](https://ai100.io/whom-to-read/charles-l-a-clarke): What systematic, replicable evaluation of question-answering systems requires — now that the systems being evaluated are language models rather than retrievers. - [Chelsea Finn](https://ai100.io/whom-to-read/chelsea-finn): How language and learning systems adapt to a new task with very few examples — and what the theoretical structure of that adaptation tells you about what they have or haven't actually learned. - [Chris Callison-Burch](https://ai100.io/whom-to-read/chris-callison-burch): Whether human readers can tell apart text produced by a language model from text produced by humans — and how that distinguishability decays as models improve. - [Christof Monz](https://ai100.io/whom-to-read/christof-monz): Year-over-year systematic evaluation of machine translation across dozens of language pairs — and what that tracking reveals about which translation problems are getting solved and which aren't. - [Christopher Manning](https://ai100.io/whom-to-read/christopher-manning): The argument that meaning, as humans use the word, is not what large language models trade in. - [Christopher Potts](https://ai100.io/whom-to-read/christopher-potts): What linguistic structure language models actually represent — and what they only seem to. - [Colin Raffel](https://ai100.io/whom-to-read/colin-raffel): Whether a single language model can do all NLP tasks at once when they're all framed as text-in-text-out — and whether that unification holds across languages. - [Dan Jurafsky](https://ai100.io/whom-to-read/dan-jurafsky): Making natural language processing a teachable discipline — and using that perspective to read where the field's current evaluation practices fit, and don't fit, into its longer history. - [Danqi Chen](https://ai100.io/whom-to-read/danqi-chen): Whether a language model that produces an answer with citations is actually grounding the answer in those citations — or just attaching plausible references after the fact. - [David Jurgens](https://ai100.io/whom-to-read/david-jurgens): Whether language models that handle factual questions cleanly can also handle the kind of social knowledge that determines what humans actually mean when they say things. - [Daxin Jiang](https://ai100.io/whom-to-read/daxin-jiang): How to pre-train an encoder so that the embeddings it produces are good for retrieval — without ever supervising on a retrieval task. - [Diyi Yang](https://ai100.io/whom-to-read/diyi-yang): How language models behave when the task is social rather than informational — persuasion, support, conflict, politeness. - [Dragan Gašević](https://ai100.io/whom-to-read/dragan-gasevic): Whether AI tools used in educational contexts actually help the learners they're built for — and what kind of evaluation infrastructure that question requires beyond conventional ML benchmarks. - [Ee-Peng Lim](https://ai100.io/whom-to-read/ee-peng-lim): Whether asking a language model to make a plan before solving a problem produces better reasoning than telling it to think step by step. - [Emma Strubell](https://ai100.io/whom-to-read/emma-strubell): Whether the compute and energy cost of training and serving language models belongs in the headline of an evaluation, where accuracy currently sits alone. - [Emmanuel Candès](https://ai100.io/whom-to-read/emmanuel-candes): How to put valid statistical uncertainty intervals around any model's predictions — including language models — without assuming you know what kind of error distribution to expect. - [Eric Horvitz](https://ai100.io/whom-to-read/eric-horvitz): The qualitative shape of language-model capability — what it looks like as a thing, and whether we have the vocabulary to describe it before we have the methodology to measure it. - [Furu Wei](https://ai100.io/whom-to-read/furu-wei): Using a language model's own parametric knowledge to make retrieval find the document the user was actually looking for. - [George J. Pappas](https://ai100.io/whom-to-read/george-j-pappas): How quickly an automated attacker can find prompts that break a language model's safety alignment — and what that means for evaluating model robustness as such. - [Gideon Mann](https://ai100.io/whom-to-read/gideon-mann): What language models look like when they're trained for a single high-value vertical — finance, in this case — and what evaluation that specialization requires beyond general-purpose benchmarks. - [Graham Neubig](https://ai100.io/whom-to-read/graham-neubig): Connecting retrieval, generation, and evaluation into a single working system — and asking when the connections actually hold. - [Hannaneh Hajishirzi](https://ai100.io/whom-to-read/hannaneh-hajishirzi): When a language model should reach for an external knowledge source — and when its own parametric memory is enough. - [Igor Mordatch](https://ai100.io/whom-to-read/igor-mordatch): What it means to evaluate a language model that takes consequential actions through its text — not only producing answers but operating in environments that respond. - [Ion Stoica](https://ai100.io/whom-to-read/ion-stoica): The distributed-compute infrastructure that almost every modern LLM is either trained on, served from, or evaluated through. - [Iryna Gurevych](https://ai100.io/whom-to-read/iryna-gurevych): Building the encoder infrastructure that makes "find sentences similar to this query" a fast and reliable operation across languages and tasks. - [James Zou](https://ai100.io/whom-to-read/james-zou): What evaluation looks like when the AI being evaluated has to clear regulatory bars before deployment — and what general LLM evaluation should learn from a field that has been doing this for years. - [Jamie Callan](https://ai100.io/whom-to-read/jamie-callan): What retrieval-augmented generation looks like when you actually know the thirty-year history of information retrieval it's reinventing. - [Jared Kaplan](https://ai100.io/whom-to-read/jared-kaplan): Whether language-model loss is a smooth function of compute, data, and model size — and what that smoothness lets you predict (and not predict) about capabilities at larger scales. - [Jennifer Wortman Vaughan](https://ai100.io/whom-to-read/jennifer-wortman-vaughan): How people actually understand and act on language-model evaluation results — and where the gap between what was measured and what gets believed. - [Jesse Dodge](https://ai100.io/whom-to-read/jesse-dodge): What has to be in a paper about a language model for another lab to be able to verify the result. - [Ji-Rong Wen](https://ai100.io/whom-to-read/ji-rong-wen): Organizing the LLM literature into something other Chinese-language NLP researchers can actually navigate from inside the academic ecosystem. - [Jian-Guang Lou](https://ai100.io/whom-to-read/jian-guang-lou): Whether a language model that produces code actually produces code that does what the request asked for — and what evaluation looks like when correctness has a definite answer for once. - [Jian-Yun Nie](https://ai100.io/whom-to-read/jian-yun-nie): How search behaves differently when the query is a conversation in a non-English language — and how language models are changing that picture. - [Jianfeng Gao](https://ai100.io/whom-to-read/jianfeng-gao): What large language models are as a class of systems — taxonomically, architecturally, and in terms of what they actually inherit from the longer history of neural NLP. - [Jie Zhou](https://ai100.io/whom-to-read/jie-zhou): Whether language models can be trusted to evaluate other language models in production NLP pipelines. - [Jimmy Lin](https://ai100.io/whom-to-read/jimmy-lin): Making information retrieval reproducible enough that an LLM researcher and an IR researcher can run the same experiment and get the same answer. - [Jimmy Xiangji Huang](https://ai100.io/whom-to-read/jimmy-xiangji-huang): How information retrieval scales over the messy operational data that real organizations hold, as opposed to the clean benchmark corpora the field actually publishes on. - [Jindong Wang](https://ai100.io/whom-to-read/jindong-wang): Standardizing what "evaluating an LLM" even means as a research procedure. - [Jochen Wirtz](https://ai100.io/whom-to-read/jochen-wirtz): What happens when customer-facing service AI replaces, augments, or competes with human service workers — and what kinds of evaluation that change actually requires. - [Jonathan Berant](https://ai100.io/whom-to-read/jonathan-berant): Whether a language model is actually doing the reasoning steps needed to answer a question — or just producing an answer that happens to be right. - [Juanzi Li](https://ai100.io/whom-to-read/juanzi-li): Combining structured knowledge from knowledge graphs with the statistical patterns language models learn from text — and what that combination buys you for evaluation. - [Julian McAuley](https://ai100.io/whom-to-read/julian-mcauley): Whether language models trained on the open web are already doing recommendation — and what that implies for products that compete with traditional recommenders. - [Junichi Yamagishi](https://ai100.io/whom-to-read/junichi-yamagishi): Whether human listeners — or automated detectors — can tell AI-generated speech apart from real human speech, and how that distinguishability changes as speech synthesis improves. - [Jure Leskovec](https://ai100.io/whom-to-read/jure-leskovec): What graph structure adds to the kinds of reasoning and retrieval problems language models currently handle without it — and what gets missed when relationships in data are flattened into text. - [Kevin Chen-Chuan Chang](https://ai100.io/whom-to-read/kevin-chen-chuan-chang): Organizing the rapidly growing literature on language-model reasoning into something a researcher new to the area can actually navigate. - [Kyle Lo](https://ai100.io/whom-to-read/kyle-lo): What's actually in a training corpus once you sit down and look at it document by document. - [Luke Zettlemoyer](https://ai100.io/whom-to-read/luke-zettlemoyer): Whether open-weight language models can be built at frontier scale and whether their factuality can be measured at fine resolution. - [Maarten de Rijke](https://ai100.io/whom-to-read/maarten-de-rijke): Whether information retrieval should be done by a system that "writes the document ID" instead of by one that searches a vector index. - [Maarten Sap](https://ai100.io/whom-to-read/maarten-sap): Whether language models can produce or reason about social knowledge with the same competence they show on factual tasks — and what's at stake when they can't. - [Maosong Sun](https://ai100.io/whom-to-read/maosong-sun): What the open-LLM ecosystem looks like when it grows out of Chinese academic NLP rather than out of Western non-profits like AI2. - [Marco Baroni](https://ai100.io/whom-to-read/marco-baroni): Whether neural language models can compose what they've learned into new combinations they've never seen — or whether they're really only doing sophisticated interpolation within their training distribution. - [Mari Ostendorf](https://ai100.io/whom-to-read/mari-ostendorf): What language technology looks like when you've been responsible for it as deployable engineering for thirty years before LLMs arrived to redo the field. - [Mark Gales](https://ai100.io/whom-to-read/mark-gales): Detecting hallucinations in a language model without any access to its weights or to ground truth. - [Martin Potthast](https://ai100.io/whom-to-read/martin-potthast): Whether language models can be used to judge whether a document is relevant to a query — and what changes when they replace human assessors in that role. - [Mengnan Du](https://ai100.io/whom-to-read/mengnan-du): What kinds of explanations can be obtained for language-model outputs — and which of those explanations turn out to be reliable. - [Michihiro Yasunaga](https://ai100.io/whom-to-read/michihiro-yasunaga): Whether a language model can correctly translate a natural-language question into a precise structured query — and what that translation reveals about reasoning over knowledge. - [Minlie Huang](https://ai100.io/whom-to-read/minlie-huang): Whether the categories of "harmful" used to evaluate language-model safety transfer from Western to Chinese-language deployment contexts, where the regulatory frame and cultural categories are different. - [Mohit Bansal](https://ai100.io/whom-to-read/mohit-bansal): Whether the methods we use to evaluate language-only models still work when the same model has to handle images, speech, or other modalities at the same time — and what fails first when modalities are combined. - [Nan Duan](https://ai100.io/whom-to-read/nan-duan): Rewriting the user's query before retrieval so that the retriever has a chance of returning useful documents. - [Nathan Lambert](https://ai100.io/whom-to-read/nathan-lambert): What happens to a language model between "trained on the internet" and "answering your question the way it does" — and how to study that step in public. - [Noah A. Smith](https://ai100.io/whom-to-read/noah-a-smith): Whether a language model can produce its own training data — and what the methodological consequences are when that becomes the standard practice. - [Norbert Fuhr](https://ai100.io/whom-to-read/norbert-fuhr): What information retrieval evaluation has been getting wrong, in writing, for the last several decades — and why each new generation of researchers makes the same mistakes. - [Omer Levy](https://ai100.io/whom-to-read/omer-levy): Whether a language model can infer the task from examples alone — without being told what to do — and what that ability reveals about how it represents instructions. - [Pang Wei Koh](https://ai100.io/whom-to-read/pang-wei-koh): What happens when a language model encounters the kind of data it didn't see during training — and how to measure that gap rigorously, not just notice it after deployment. - [Paolo Rosso](https://ai100.io/whom-to-read/paolo-rosso): Building evaluation campaigns that work for Iberian-Romance languages — Spanish, Catalan, Portuguese — instead of porting English-centric methodology and accepting the resulting blind spots. - [Pascale Fung](https://ai100.io/whom-to-read/pascale-fung): Whether the same language model performs the same kind of work across different languages — and whether evaluation methodology that's built for English is misleading us about what models do in everything else. - [Percy Liang](https://ai100.io/whom-to-read/percy-liang): How to measure language models so a measurement made today still means something next year. - [Peter Henderson](https://ai100.io/whom-to-read/peter-henderson): Where the technical findings about language-model behavior actually matter — in audits, regulation, and legal liability — and what the gap between "we measured this" and "this changes what's allowed" looks like. - [Philip S. Yu](https://ai100.io/whom-to-read/philip-s-yu): Bridging four decades of data-mining methodology to the question of how to evaluate large language models without reinventing techniques the field already has. - [Prateek Mittal](https://ai100.io/whom-to-read/prateek-mittal): How visual inputs become a new attack surface for safety-aligned language models that accept multimodal queries. - [Qiang Yang](https://ai100.io/whom-to-read/qiang-yang): Whether useful machine learning can happen when the data you'd train or evaluate on can't be moved to a single place — for legal, privacy, or commercial reasons. - [Quoc V. Le](https://ai100.io/whom-to-read/quoc-v-le): The architectural and training-recipe building blocks that the modern LLM era was built on top of. - [Rishi Bommasani](https://ai100.io/whom-to-read/rishi-bommasani): What language-model developers are and aren't telling us about their own models — and how to measure that systematically. - [Roi Reichart](https://ai100.io/whom-to-read/roi-reichart): What "domain" means for a language model — when its training distribution stops matching its deployment context — and how to evaluate that mismatch rigorously. - [Seungone Kim](https://ai100.io/whom-to-read/seungone-kim): Whether the field can have an open-weight LLM-evaluator alternative, so that "one black box judging another black box" stops being the only available option. - [Shafiq Joty](https://ai100.io/whom-to-read/shafiq-joty): Evaluation methodology for language models when the deployment context is enterprise software rather than a research demo. - [Shayne Longpre](https://ai100.io/whom-to-read/shayne-longpre): What's actually inside the data language models train on — and who can or can't tell. - [Shinji Watanabe](https://ai100.io/whom-to-read/shinji-watanabe): The open-source speech-processing infrastructure that lets academic and industrial groups train, evaluate, and compare voice-input or voice-output language systems on the same footing. - [Shuming Shi](https://ai100.io/whom-to-read/shuming-shi): What "language model hallucination" looks like when you're responsible for shipping LLM-powered products to a billion-user surface. - [Steven Schockaert](https://ai100.io/whom-to-read/steven-schockaert): How to evaluate a retrieval-augmented system end-to-end when each part of it can fail in different ways. - [Tatsunori Hashimoto](https://ai100.io/whom-to-read/tatsunori-hashimoto): Whether the numbers reported about language models are statistical findings or methodological artifacts. - [Tom Mitchell](https://ai100.io/whom-to-read/tom-mitchell): Whether a language model has an internal representation of whether it's telling the truth — separable from what it actually outputs. - [Torsten Hoefler](https://ai100.io/whom-to-read/torsten-hoefler): Generalizing language-model reasoning beyond linear chains of thought — into branching, backtracking, and recombination of intermediate reasoning steps. - [Tushar Khot](https://ai100.io/whom-to-read/tushar-khot): What "reasoning ability" actually means as something you can put on a benchmark — and what changes when you also try to coach the model to do reasoning through structured prompting. - [Wayne Xin Zhao](https://ai100.io/whom-to-read/wayne-xin-zhao): The synthesis side of LLM research — what it takes to read every paper of the moment and produce something other researchers can navigate. - [Weijia Shi](https://ai100.io/whom-to-read/weijia-shi): Making retrieval-augmented generation work when the language model itself is a closed box you can't fine-tune. - [Wen-tau Yih](https://ai100.io/whom-to-read/wen-tau-yih): Making the retrieval step inside retrieval-augmented systems good enough that the generation step has something to work with. - [Wenjie Li](https://ai100.io/whom-to-read/wenjie-li): How generative retrieval relates to the other generation tasks — summarization, question answering — that the same system architecture has to handle. - [Xia Hu](https://ai100.io/whom-to-read/xia-hu): Whether the academic state of the art in language models can be turned into something a practitioner — not a research lab — can actually deploy and trust. - [Xing Xie](https://ai100.io/whom-to-read/xing-xie): Connecting the recommender-systems tradition of measuring user-facing AI behavior with the new evaluation challenges modern LLMs pose. - [Xipeng Qiu](https://ai100.io/whom-to-read/xipeng-qiu): Building an open Chinese-language large language model that the academic community can actually study under the hood — and the tooling around it. - [Xueqi Cheng](https://ai100.io/whom-to-read/xueqi-cheng): What information retrieval research from inside the Chinese IR tradition has been arguing — and how its angle on generative retrieval differs from the Western canon. - [Yang Liu](https://ai100.io/whom-to-read/yang-liu): Whether a more capable language model can be used to grade the outputs of a less capable one — and how to do that without fooling yourself. - [Yann LeCun](https://ai100.io/whom-to-read/yann-lecun): What machine-learning systems should look like, set against what they currently are. - [Yarin Gal](https://ai100.io/whom-to-read/yarin-gal): Quantifying when language models don't know what they're saying. - [Yejin Choi](https://ai100.io/whom-to-read/yejin-choi): What it would take for a language model to have what humans have in spades and what LLMs reliably lack — common sense. - [Yoav Goldberg](https://ai100.io/whom-to-read/yoav-goldberg): Whether the chain of reasoning a language model produces is the chain of reasoning it actually followed. - [Yoav Shoham](https://ai100.io/whom-to-read/yoav-shoham): What can be measured about the state of AI from outside any single company — and what retrieval-augmentation looks like when you don't need to retrain anything. - [Yonatan Belinkov](https://ai100.io/whom-to-read/yonatan-belinkov): What's actually inside a language model's hidden representations — and which of those internal states map onto things humans would recognize as knowledge. - [Yue Zhang](https://ai100.io/whom-to-read/yue-zhang): Organizing what the field calls "hallucination" into categories that actually mean different things. - [Yulia Tsvetkov](https://ai100.io/whom-to-read/yulia-tsvetkov): What kinds of bias and harm look like in language-model outputs across languages — especially the languages and communities the field's standard evaluation has historically ignored. - [Zhaochun Ren](https://ai100.io/whom-to-read/zhaochun-ren): Whether a language model can be trusted with the job that's currently done by an information retrieval system — and which parts of that job it actually does well. - [Zhiting Hu](https://ai100.io/whom-to-read/zhiting-hu): Treating a language model as one component in a larger planning system — instead of asking it to do reasoning end-to-end in its own head. ## Legal - [Terms of Service](https://ai100.io/terms) - [Privacy Policy](https://ai100.io/privacy) - [Refund Policy](https://ai100.io/refund) - [Delivery Policy](https://ai100.io/delivery) - [Imprint](https://ai100.io/imprint)