Overview

Arena is a community platform designed to use, compare, and evaluate artificial intelligence models in real-world situations. It allows the same request to be sent to several models, their results to be compared, and a preferred answer to be selected.

The project was launched in 2023 by researchers from UC Berkeley and the LMSYS ecosystem under the name Chatbot Arena. It was initially an academic experiment dedicated to large language models and evaluation through human preferences.

Chatbot Arena later became LMArena, then simply Arena in January 2026. This name change reflects the project’s expansion beyond language models alone: the platform now covers text, code, Web search, vision, documents, images, videos, interface development, and agents.

Arena is now operated by Arena Intelligence, Inc. The company preserves the community-based and scientific approach of the original project while developing commercial evaluation services for artificial intelligence laboratories, companies, and developers.

Arena’s historical operation is based on Battle Mode. Two anonymous models receive the same prompt and each produce an answer. The user compares the results without knowing the models’ identities, then votes for the better answer, declares a tie, or indicates that both results are insufficient.

The model names are revealed after the vote. This anonymisation aims to reduce the influence of brand, reputation, or user expectations on the judgement.

Votes are then aggregated to construct public rankings. Unlike a benchmark composed of a fixed list of questions, Arena relies on real requests submitted by its community. Models are therefore confronted with a very wide range of uses, languages, styles, and difficulty levels.

The platform is no longer limited to anonymous battles. Direct Mode allows users to select a specific model and use it as a conventional assistant. Side by Side allows two identified models to be selected for direct comparison.

Arena also provides Max, a router that analyses the request and automatically selects a model according to its observed performance and latency. Max covers several modalities, including text, search, vision, image generation and editing, and front-end development.

Code Arena compares models on the creation of interfaces, websites, and applications. Users can inspect the proposals, preview the results, examine the code, and continue the conversation to improve both versions.

Agent Mode extends the platform to longer and more structured tasks. The agent can search for information, manipulate files, produce images, write code, and use a terminal-like or sandbox environment.

The various votes feed several specialised rankings. A model may therefore perform well in general writing without holding the same position for code, search, document analysis, image generation, or video.

Arena is therefore simultaneously a free tool for trying many models, an interactive comparison service, a platform for collecting human preferences, and a public infrastructure for evaluating artificial intelligence.

The Web service remains proprietary. Arena nevertheless publishes some of its data, research, and methodological tools. The Arena-Rank engine used to calculate the rankings is notably available under the Apache-2.0 licence.

Features

  • Battle Mode: compare two models receiving exactly the same request.

  • Anonymous models: hide their identities during comparison to reduce brand-related bias.

  • Reveal after voting: display the names of both models once the judgement has been recorded.

  • Vote for answer A or B: select the result considered the most useful or successful.

  • Tie: indicate that both answers provide comparable quality.

  • Both answers insufficient: report that both models failed or produced an unusable result.

  • Multi-turn conversations: continue an exchange to compare how models retain context.

  • Direct Mode: use an explicitly selected model without an anonymous confrontation.

  • Side by Side: select two identified models for direct comparison.

  • Model selector: list the available models and sort or filter them according to their modalities.

  • Modality information: icons showing whether a model supports text, vision, images, video, or other capabilities.

  • Proprietary models: access commercial models supplied by their publishers.

  • Open-weight models: access models whose weights can be downloaded under different licences.

  • Early access to new models: occasionally evaluate models before their definitive public launch.

  • Max Smart Router: automatically select the model considered suitable for a request.

  • Performance-based routing: use community results to guide model selection.

  • Latency awareness: seek a balance between result quality and response speed.

  • Multimodal Max: provide routing for several categories beyond text.

  • Web search with Max: automatically select a model or system suited to questions requiring online information.

  • Vision with Max: analyse requests containing images.

  • Image generation with Max: automatically route the request to a suitable visual model.

  • Image editing with Max: choose a model specialising in visual transformation.

  • Front-end development with Max: route requests for creating pages or interfaces.

  • Text Arena: compare models on conversation, writing, reasoning, and general-purpose tasks.

  • Code Arena: compare models capable of producing code and interfaces.

  • WebDev Arena: ranking dedicated to the creation of applications and websites.

  • Image-to-WebDev Arena: evaluate the ability to transform a visual reference into a Web interface.

  • Vision Arena: compare models capable of analysing images.

  • Document Arena: query and compare models using imported documents, including PDFs.

  • Search Arena: compare systems using Web search to produce documented answers.

  • Text-to-Image Arena: compare models generating images from a prompt.

  • Image Edit Arena: compare models that modify or transform images.

  • Text-to-Video Arena: evaluate models generating video from a description.

  • Image-to-Video Arena: animate or transform an image into video.

  • Video Edit Arena: compare systems capable of modifying videos.

  • Agent Arena: evaluate models or orchestrators performing complex tasks with several tools.

  • Image Arena: generate two anonymous images from the same prompt.

  • Visual comparison: choose the most faithful, attractive, or relevant result according to the request.

  • Video Arena: generate and compare videos from several models.

  • File uploads: add compatible documents or resources in selected modes.

  • PDF analysis: ask questions, request summaries, and extract information from documents.

  • Image analysis: describe, interpret, or reason from a visual input.

  • Web search: produce answers using recent information and online sources.

  • Image generation: limited free use of visual-generation models.

  • Image editing: transform visual content supplied by the user.

  • Video generation: create sequences from text or images depending on the available models.

  • Agent Mode: perform workflows that are longer than conventional interactions.

  • Search integrated into agents: collect information from the Web during a task.

  • Programming tools: generate and modify code within an agentic workflow.

  • Sandbox and Bash: execution environment allowing an agent to perform certain technical operations.

  • File creation: produce resources during a complex task.

  • Context retention: continue a project across several stages of the same exchange.

  • Agent evaluation: ranking based on several signals extending beyond a single final vote.

  • Confirmed Success: measure the confirmed completion of an agentic task.

  • Praise vs Complaint: take positive or negative user reactions into account.

  • Steerability: evaluate an agent’s ability to follow corrections and new instructions.

  • Bash Recovery: analyse how the agent recovers after an execution error.

  • Tool Hallucination: detect tools or actions claimed by the agent without actually being used.

  • Interface previews: directly display applications produced in Code Arena.

  • Code display: inspect files or blocks generated by the models.

  • Conversational iterations: progressively improve both proposals in Code Arena.

  • Preview links: share certain generated prototypes.

  • Public leaderboards: consult rankings without creating an account.

  • General ranking: overview of the main results across several arenas.

  • Specialised rankings: separate models according to their task or modality.

  • Rank: the model’s raw position according to its central score.

  • Rank Spread: the range of possible positions given the statistical uncertainty.

  • Arena Score: relative estimate of model performance based on its battles.

  • Confidence interval: indication of the uncertainty associated with the score.

  • Number of votes: volume of comparisons used to evaluate the model.

  • Price per million tokens: indicative input and output pricing information when available.

  • Dynamic updates: scores change as new votes are collected.

  • Model filters: select according to provider, category, or other characteristics.

  • Expert Leaderboard: ranking based on prompts submitted or selected by expert contributors.

  • Occupational Leaderboards: analyse performance according to professional fields.

  • Profession-based rankings: compare models in categories such as development, research, law, or other available sectors.

  • Factuality Leaderboard: ranking combining human preference with the accuracy of claims.

  • Factuality signal: automated verification of claims that can be checked against Web sources.

  • Configurable factuality weighting: assign a specific weight to the factual-correctness signal.

  • Text ranking with factuality: compare text models while accounting for accuracy.

  • Search ranking with factuality: evaluate search models according to response quality and correctness.

  • Style control: methodology designed to limit artificial advantages linked to particular presentation habits.

  • Bradley-Terry model: estimate relative performance from pairwise battle results.

  • Arena-Rank: Python package used to calculate scores and rankings.

  • Statistically calculated intervals: quantify uncertainty instead of presenting a position as an absolute truth.

  • Model rebalancing: method limiting certain effects caused by differences in battle volume.

  • Open-source methodology: publication of the ranking engine under the Apache-2.0 licence.

  • Partial public data: release selected conversations, votes, or results as datasets.

  • Leaderboard history: provide data for studying how rankings evolve.

  • Methodological changelog: record changes made to the evaluation system.

  • Scientific publications: document methods and results in several research papers.

  • Leaderboard inclusion criteria: require models to be publicly available or comply with preview rules.

  • Eligible open-weight models: allow inclusion when a model can be downloaded publicly.

  • Models available through an API: allow models with a public and documented API.

  • Eligible public services: include models available within a widely accessible application.

  • Pre-launch evaluation: provide a controlled process for testing a model before its public release.

  • Minimum number of battles: collect enough votes before treating a score as stable.

  • Account creation by email: register with a verified email address.

  • Google sign-in: authenticate with a Google account.

  • Conversation history: retain chats linked to the account.

  • Cross-device synchronisation: retrieve history after signing in from another browser.

  • Conversation archiving: organise saved exchanges.

  • Session deletion: remove selected chats from the user interface.

  • Leaderboard access without an account: publicly consult the rankings.

  • Additional features after signing in: access Direct, Side by Side, Image Arena, and Video Arena.

  • Higher limits for signed-in accounts: provide different quotas for anonymous and registered users.

  • Model-specific limits: restrict requests to particularly costly or popular models.

  • Global chat limit: cap the total volume of usage.

  • Daily usage balance: limiting mechanism designed to distribute available resources.

  • Reset after twenty-four hours: automatically restore access after the daily quota has been exhausted.

  • Free access without a paid priority option: no current public subscription that removes the limits.

  • Commercial evaluations: services intended for model providers and companies.

  • Reports based on human preferences: professional analysis of model performance on real-world uses.

Use cases

Comparing two models without knowing their brand

Battle Mode sends the same question to two anonymous models and allows their answers to be compared directly.

The user judges only the visible result, without being influenced by the names OpenAI, Anthropic, Google, Meta, Mistral, or another provider.

This approach is useful for determining which model genuinely performs best for a specific request instead of relying solely on its reputation.

Choosing a general-purpose assistant

Arena can be used to test several models on common assistant tasks: writing, explanation, summarisation, translation, analysis, or idea generation.

By repeating comparisons with their own prompts, users can identify models whose tone, accuracy, and level of detail best match their expectations.

The general ranking provides an initial indication, but personal testing remains more relevant when choosing an everyday tool.

Checking which model writes best

The same writing instruction can be submitted to two models to compare structure, style, creativity, and compliance with constraints.

This method suits articles, descriptions, scripts, dialogue, summaries, marketing content, and social posts.

The vote remains subjective: a more elegant answer is not necessarily more accurate or better suited to a particular editorial direction.

Comparing models for programming

Code Arena allows several models to be compared on the creation of a component, interface, or application.

Both results can be previewed and their code examined before voting.

A developer can therefore evaluate not only the final appearance, but also the readability, structure, maintainability, and technical compliance of the code.

Rapidly building a Web prototype

Users can describe a page, application, or component and allow two models to produce their proposals.

The conversation can then continue to add features, correct the interface, or modify the design.

This approach suits mock-ups, demonstrations, landing pages, and early product exploration.

Comparing a reproduction from an image

Image-to-WebDev allows a visual reference to be supplied and shows how several models transform it into an interface.

This feature can test fidelity to framing, colours, typography, and visual hierarchy.

The result must nevertheless be reviewed for accessibility, responsive design, and code quality.

Choosing an image-generation model

Image Arena sends the same prompt to two anonymous models and displays their results side by side.

Users can compare compliance with the request, composition, style, details, anatomy, and the quality of integrated text.

Votes then contribute to the Text-to-Image ranking.

Evaluating a generative-editing tool

Image Edit Arena compares how several models transform the same image according to an instruction.

The evaluation may focus on identity preservation, background consistency, accuracy of modifications, and the absence of artefacts.

It is useful for selecting a model for object replacement, style changes, or local modifications.

Comparing video generators

Video Arena compares text-to-video or image-to-video models.

Criteria can include motion consistency, prompt adherence, character stability, visual quality, and temporal continuity.

The ranking can help identify the most popular models, but every video should still be viewed in full.

Testing visual understanding

Vision Arena compares answers from models receiving an image and a question.

This feature can describe a scene, read an interface, interpret a chart, or analyse a creative reference.

Recognition and interpretation errors remain possible, particularly when the visual is ambiguous or highly detailed.

Querying a document

Document Arena allows users to upload a document, including a PDF, and then ask questions of several models.

Users can compare their ability to summarise, retrieve information, interpret a table, or follow an instruction related to the content.

This feature should not be used with confidential documents or files containing sensitive information.

Comparing Web-research assistants

Search Arena evaluates models capable of searching for recent information and producing a documented answer.

Users can examine source relevance, subject coverage, citation quality, and the final synthesis.

The factuality filter provides an additional signal regarding the accuracy of claims without replacing human verification.

Identifying a model by profession

Occupational rankings show differences between models across work-related categories.

A model that ranks highly overall may perform less effectively on specialised legal, scientific, technical, or writing requests.

These filters help avoid reducing model quality to a single general score.

Examining performance on expert prompts

The Expert Leaderboard focuses on requests submitted or selected by contributors with specific expertise.

It can provide a more demanding signal than general-purpose community prompts.

The representation of disciplines and expert profiles should nevertheless be considered when interpreting the ranking.

Testing a model before subscribing to its service

Arena provides access to many proprietary models without requiring a separate subscription from each provider.

Users can try an answer, compare several models, and identify those that deserve more extensive testing.

The quotas, parameters, and integrations provided by Arena do not necessarily reproduce the complete experience of the official application.

Letting Max select the model

Max suits users who do not want to choose manually between dozens of models.

The router analyses the modality and request, then selects a model while accounting for performance and latency.

This simplicity reduces control over the routing decision and does not guarantee that the selected model will always be optimal for a highly specific need.

Performing a complex task with an agent

Agent Mode can combine research, analysis, programming, image generation, and file manipulation within the same workflow.

It suits project preparation, prototype creation, in-depth research, or the organisation of several successive operations.

The tools used and commands executed should be monitored, particularly when files or a Bash environment are involved.

Tracking changes in the AI market

The leaderboards make it possible to observe new versions, ranking changes, and the arrival of new providers.

They are useful resources for monitoring text, code, image, video, search, and agent models.

A position should always be interpreted alongside its date, number of votes, and confidence interval.

Studying user preferences

Arena data and publications can be used to analyse how humans evaluate model answers.

Researchers can study stylistic biases, language differences, professional fields, or changes in preferences.

Only part of the data is made public, after various processing and filtering steps.

Reproducing a ranking system

Arena-Rank allows researchers and developers to calculate their own rankings from battle data.

The package can be applied to artificial intelligence models, but also to other competitions based on pairwise comparisons.

Building a reliable system nevertheless requires appropriate sampling, data cleaning, and statistical interpretation.

PANACHES review

Arena occupies a distinctive place in the artificial intelligence ecosystem. It is neither a simple assistant, a conventional aggregator, nor a fixed benchmark: the platform combines model access, interactive comparison, preference collection, and methodological research.

Its main strength lies in anonymous comparison. Users often know a model’s reputation before reading its answer. Hiding its identity limits part of this bias and forces the user to judge the result that was actually produced.

This anonymisation does not eliminate every bias. Certain writing signatures, citation formats, or stylistic habits may allow a provider to be recognised. Judgement also remains influenced by the user’s personal preferences.

The collection of real prompts is another important advantage. Traditional benchmarks eventually become known, optimised for, or contaminated by training data. Arena continually confronts models with new requests drawn from everyday uses.

This diversity brings the evaluation closer to real-world conditions, but makes the data more difficult to control. Prompts may be imprecise, subjective, repetitive, or concentrated on topics popular within the community.

The ranking should therefore not be interpreted as an absolute measurement of intelligence. It reflects the average preference observed across a collection of battles, during a particular period, and with the models available on the platform.

The Rank Spread is often more informative than the raw rank. Two models displayed in third and sixth place may be statistically difficult to distinguish when their intervals overlap.

The number of votes also matters. A recently added model may display a high score with considerably more uncertainty than a model evaluated across tens of thousands of battles.

The separation between arenas is an essential improvement. The best general-purpose model is not necessarily the best for creating an interface, editing an image, analysing a PDF, or producing a video.

Professional and expert filters reinforce this specialised interpretation. They make it possible to search for performance relevant to a profession rather than blindly following the global ranking.

Adding a factuality signal addresses an important weakness of human voting. A long, confident, and well-presented answer may be preferred even when it contains several errors.

The factuality ranking does not fully solve the problem. Automated verification depends on extracting claims, searching the Web, the sources available, and the system responsible for deciding whether they are correct.

Arena nevertheless remains more transparent than many commercial rankings. The publication of Arena-Rank, some data, scientific papers, and the changelog allows a significant part of the process to be examined.

The platform is also becoming a genuine model aggregator. Direct Mode and Side by Side provide free access to models that would normally require several separate accounts or subscriptions.

This access is particularly useful for discovering a new model. Users should nevertheless avoid assuming too quickly that its behaviour within Arena will be identical to that of its official application.

Providers may use different system prompts, settings, quotas, or endpoints. Certain native features, memories, connectors, projects, or advanced parameters are not reproduced.

Max represents an interesting approach to automatic selection. Instead of requiring users to understand the strengths of every model, Arena uses its evaluation data to route the request.

This idea directly aligns with PANACHES’ needs around intelligent model selection. A coding task, document search, image generation, or lightweight conversation should not necessarily use the same engine.

Automatic routing nevertheless introduces another opaque layer. For reproducible projects, explicitly fixing the model may be preferable to allowing the router to evolve with its updates.

Agent Mode also demonstrates Arena’s development towards a more complete working environment. The platform no longer measures only an isolated answer, but the ability to complete a task, use tools, and recover from an error.

This evaluation is closer to modern agent use, but also becomes more complex. Success depends on the model, orchestrator, tools, sandbox, permissions, and the way the user formulates feedback.

The Image and Video arenas make Arena particularly important for creators. They reveal differences between models whose official demonstrations are often difficult to compare objectively.

Visual voting nevertheless remains highly subjective. A spectacular image may be preferred to one that follows the prompt more accurately, while a dynamic video may conceal continuity problems.

For PANACHES Media, Arena is an excellent monitoring source. The rankings make it possible to track new models and rapidly identify tools improving in text, code, images, video, or search.

Positions should not be quoted alone in an article. A serious analysis should specify the arena, date, number of votes, uncertainty, method, and limitations of the signal being used.

The main concern relates to confidentiality. Prompts, images, documents, votes, and interactions are precisely the data Arena needs to improve its evaluations.

The privacy policy states that some data may be used for research, service improvement, and product development. It may also be transmitted to the providers of the models being used.

The terms also grant Arena a very broad licence over submitted and generated content. This makes the service unsuitable for trade secrets, unpublished work, personal data, and confidential projects.

For PANACHES, Arena should therefore be viewed as a public comparison laboratory rather than a private space for internal documents.

Its entirely remote nature also moves it away from PANACHES’ local-first philosophy. Models, quotas, availability, and policies may change without local control.

Arena nevertheless remains an essential reference. Few platforms combine such a large community, so many modalities, direct model access, and a partially auditable methodology.

The tool therefore deserves a central place in the PANACHES directory, both as a comparison service, monitoring source, experimentation environment, and example of routing based on real evaluations.

Points to consider

  • The official name is Arena: LMArena and Chatbot Arena are former names that still appear in certain URLs, publications, or third-party pages.

  • Do not confuse Arena with a simple benchmark list: the service also allows direct use of models and agents.

  • Rankings change continuously: a position observed today may change after new votes or a methodological update.

  • Always record the date: a screenshot or article should specify when the leaderboard was consulted.

  • Identify the relevant arena: Text, Code, Search, Vision, Document, Image, and Video do not measure the same capabilities.

  • Do not rely only on the general ranking: a model can excel in one category and remain average in another.

  • Examine the Rank Spread: the raw rank may exaggerate the difference between statistically close models.

  • Read the confidence intervals: a central score is only an estimate accompanied by uncertainty.

  • Check the number of votes: a recent model may have a much smaller sample.

  • Rankings are relative: the score depends on the other models included in the battles.

  • The results do not measure every quality: safety, cost, energy use, privacy, and ease of deployment are not summarised by the score.

  • Votes reflect preference: they do not guarantee truth, safety, or professional suitability.

  • Style can influence the vote: a longer, more structured, or more confident answer may be preferred regardless of its accuracy.

  • Factuality remains partial: not every claim can be automatically verified on the Web.

  • The factuality signal is weighted: it does not entirely replace human preferences in the relevant rankings.

  • Prompts are self-selected by the community: the distribution of uses does not necessarily match that of a company or country.

  • Language biases are possible: the most widely used languages may receive more votes and comparisons.

  • Demographic biases are possible: Arena users do not necessarily represent the entire population.

  • Models may still be recognisable: certain formulations or presentation styles may indirectly reveal their identities.

  • A vote does not replace an evaluation framework: defining custom criteria remains preferable for professional decisions.

  • Test with your own tasks: project-specific prompts are often more useful than the global position.

  • Compare several times: a single generation may be affected by model randomness.

  • Models may use specific settings: their configuration within Arena may differ from that of the official application.

  • Certain native features are absent: memory, projects, connectors, tools, or personalisation may not be available.

  • Model availability varies: a provider may restrict, replace, or remove access.

  • Previews are temporary: a model tested before release may change or disappear from the ranking.

  • Check identity between preview and public release: Arena provides rules, but differences remain a point to monitor.

  • The price per million tokens is indicative: it generally describes the provider’s API price, not the cost of using Arena.

  • Free access is limited: free does not mean unlimited.

  • Daily limits apply: usage may be blocked until the quota resets.

  • Limits vary by model: a popular or expensive model may become unavailable before others.

  • Global chat limits apply: several models may be affected by the same general cap.

  • No public priority plan exists: Arena currently offers no payment option to remove these restrictions.

  • Quotas may change: the daily system remains subject to modification.

  • Creating an account increases access: Direct, Side by Side, Image, and Video generally require sign-in.

  • History depends on the account: anonymous conversations may not be recoverable.

  • One account identity: the terms notably restrict the multiplication of accounts.

  • Adults only: the terms state a minimum age of eighteen.

  • Personal or internal use: the terms generally restrict use to these contexts unless an additional agreement applies.

  • Unauthorised automation is prohibited: programmatic queries, bots, and automated extraction are restricted by the terms.

  • Do not scrape the rankings: automated extraction of names, scores, and data may violate the service rules.

  • Do not manipulate votes: fake votes, coordinated campaigns, and compensation in exchange for votes are prohibited.

  • Vote honestly: systematically choosing one side or attempting to identify a brand degrades the quality of the signal.

  • Do not send confidential data: prompts and files may be processed by Arena and third-party providers.

  • Do not upload trade secrets: private code, contracts, strategies, and unpublished documents should remain outside the platform.

  • Avoid personal data: names, contact details, medical records, financial data, and sensitive identifiers should not be transmitted.

  • Model providers may receive the content: data required for generation is sent to the relevant third-party services.

  • Providers may apply their own policies: each model may be governed by different terms and practices.

  • Prompts may be used for research: Arena states that votes and requests may be used to produce studies and analyses.

  • Data may contribute to improvement: content may be used to develop Arena or related technologies.

  • Some data may be published: conversations, inputs, and results may appear in datasets or publications after processing.

  • Do not submit anything that should never become public: the policy explicitly recommends avoiding sensitive content.

  • Broad licence over content: the terms grant Arena worldwide, transferable, perpetual, and sublicensable rights.

  • Content remains attributed to the user between the parties: this contractual ownership remains subject to the licence granted to the service.

  • Outputs may not receive copyright protection: Arena does not guarantee their eligibility for legal protection.

  • Outputs are not necessarily unique: other users may obtain similar results.

  • Respect rights over uploaded files: the user must hold the necessary permissions.

  • Review images of people: using a recognisable likeness may involve image rights.

  • Review every answer: models may produce false, outdated, or invented information.

  • Verify citations: Search Arena may display a source that does not precisely support the associated claim.

  • Do not use it alone for sensitive decisions: health, law, finance, and safety require appropriate sources and professionals.

  • Inspect files produced by agents: an apparently successful task may contain hidden errors.

  • Monitor Bash commands: a command may modify, delete, or expose files in the environment.

  • Check which tools were actually used: tool hallucinations are among the problems evaluated in Agent Arena.

  • Review generated code: a visually successful interface may contain unnecessary dependencies or vulnerabilities.

  • Test responsive design: a preview that works on one screen does not guarantee mobile compatibility.

  • Check accessibility: contrast, keyboard navigation, labels, and semantic structure must be reviewed.

  • Inspect generated images: anatomy, text, logos, and repetitive details may contain defects.

  • View videos in full: artefacts, identity transformations, and inconsistencies may appear during movement.

  • Compare fidelity, not only appearance: the most spectacular result is not always the one that best follows the instruction.

  • Arena-Rank is not the entire platform: opening the ranking engine does not make the complete Web service open source.

  • FastChat is a separate historical project: it formed part of the foundation for Chatbot Arena but is no longer the complete code for the current platform.

  • Multiple licences apply: Arena-Rank, datasets, historical tools, and third-party models may use different licences.

  • Check each model’s licence: open weight does not necessarily mean free of commercial restrictions.

  • Public datasets are partial: they do not necessarily reproduce the complete real-time leaderboard.

  • Methods may change: weighting, filters, intervals, and inclusion rules evolve with the research.

  • Consult the changelog: a ranking variation may result from a methodological update rather than a model change.

  • Keep your own results: important notes, screenshots, and prompts should be archived outside the service.

  • Plan an alternative: free access to a model may be reduced or interrupted.

  • Do not rely on Arena as an API: the platform is not designed as an automated endpoint for an application.

  • Use the provider’s official API for production: it generally offers more suitable contracts, quotas, and parameters.

  • Compare with complementary evaluations: automated benchmarks, internal tests, costs, and technical constraints should supplement Arena.

  • Do not confuse popularity with personal relevance: the best average model is not necessarily the best one for the user’s workflow.

  • Choose by task: text, code, search, images, video, and agents may justify different models.

  • Treat Arena as a public laboratory: its strength lies in exploration and comparison, not confidentiality.