Table of Contents
- Key Takeaways
- 1. AI Expanding Both Inside and Outside the Exchange App
- 2. Exchanges and AI: The Execution Layer and the Decision Layer
- 2.1 Exchange AI Features Viewed Through the Two Layers
- 2.2 The More Execution Is Opened Up, the More Users Move Outside
- 3. The Difficulty of Building a Decision Layer and Lessons from Traditional Finance
- 3.1 Why Building a Decision Layer Is Difficult
- 3.2 How Traditional Finance Built Compliant LLM Services
- 4. Building a Decision Layer: The Case of Bithumb and Surf
- 4.2 Surf's Four Design Choices for Keeping the Model in Check
- 4.3 What Sets Bithumb AI Apart
- 5. Exchange Intelligence Is Completed Outside the Model
Researcher
Related Projects
Key Takeaways
- The use of AI at CEXs can be divided into two layers: an execution layer that turns user intent into orders, and a decision layer that produces the information and interpretation needed to decide what to buy and sell. The decision layer splits further into a conversational type, which returns the model's answers as they are, and a publishing type, in which the exchange sets the evidence and verification standards and releases the results under its own name. The more an exchange opens execution to outside agents, the more user judgment and trust are formed outside the exchange. This is where the incentive for exchanges to build their own publishing-type decision layer comes from.
- Traditional finance was able to deploy language models within compliance not because it trusted the models, but because it built systems that did not require trusting them. For a language model to be used in finance while remaining compliant, four measures are needed: selecting the material the model may reference, defining what the model will not do, evaluating outputs against expert standards, and placing a final checkpoint before anything is released.
- Bithumb AI, jointly developed by Bithumb and Surf, implements these four measures through point-in-time snapshots, the separation of calculation from explanation, an evaluator trained on reviewer labels, and the withholding of publication. The case of Bithumb AI shows that the competitiveness of exchange intelligence lies not in the type of analysis offered, but in the standards that govern everything from data selection to publication.
1. AI Expanding Both Inside and Outside the Exchange App
Exchanges have long competed on factors such as listed assets, fees, liquidity, derivatives, and regulatory licenses. AI is now being added as a new element of competition. The scope of that competition is also widening, from analysis features inside the exchange app to integration with external AI agents.
AI inside exchange apps is not new in itself. A representative example is Bybit, which launched TradeGPT, a ChatGPT-based tool, in 2023 to provide market analysis and Q&A within its app. Features of this kind helped users find and interpret information, but they functioned as little more than a chatbot placed next to charts and news, and did not provide information substantial enough to play a direct part in users' decisions.
This year, driven by improvements in model performance and a broader range of uses, the ways in which AI agents outside the exchange can access exchange functions have also expanded. In March, OKX released the Agent Trade Kit, which uses the Model Context Protocol (MCP). It is a tool that allows agents running on external models such as Claude to look up prices and execute orders. In the same month, Bybit and Binance also released tools that support external agents' access to their exchanges. In June, Coinbase made order and payment functions available to agents through Coinbase for Agents. In Korea, Upbit released "Upbit Skills" in May, which teaches AI coding agents how to use its API. Bithumb followed in June with the AI Trade Kit, which lets users place orders through conversations with ChatGPT or Claude.
This is what makes the current competition different from before. Listings, fees, and liquidity are all variables that an exchange can adjust at its own discretion, so listing faster, cutting fees, and attracting liquidity were enough to get ahead. On the AI axis, by contrast, the value users receive comes from the ability to read the market and decide what to buy and sell. Under the current structure, however, that ability belongs to external models such as ChatGPT or Claude, and the exchange merely opens an account so the model can place orders. In effect, the core of the competition that exchanges have rushed into lies somewhere the exchanges themselves can neither build nor fix.
Against this backdrop, what AI can do at an exchange falls broadly into two layers.
- Execution layer: turning user intent into orders
- Decision layer: producing the information and interpretation needed to decide what to buy and sell
The decision layer is further divided into two types, depending on where responsibility lies. The first is the conversational type, in which the user formulates the question and the model's answer is returned without separate verification by the exchange. Chatbot-style tools belong here.
The second is the publishing type, which covers features where the exchange sets the format, update cycle, evidence, and verification procedures, and then releases the results under its own name. Publishing at the level of news summaries already existed, as with Upbit's AI news briefing. Attempts to go beyond market summaries and publish judgments on individual assets or a user's portfolio under the exchange's name, however, have only emerged recently. Examples include Coinbase Advisor, launched by Coinbase in June, and Bithumb AI, jointly developed with Surf and launched by Bithumb in September.
In this article, I first examine how functions in the execution layer and the publishing-type decision layer place different responsibilities on exchanges. I then look at why building a decision layer is difficult and how traditional finance has addressed that problem. Finally, through the case of Bithumb and Surf, I discuss how a crypto exchange can build a decision layer and what questions remain.
2. Exchanges and AI: The Execution Layer and the Decision Layer
The biggest difference between the two layers lies less in their functions than in the responsibility the exchange bears. In the execution layer, the exchange must answer two questions: what to open to external models and how far, and whether incoming orders were processed accurately. In the decision layer, the focus shifts to whether the explanations sent to users are correct, how confidently they were stated, and whether errors were filtered out. Questions in the execution layer can be answered through design outside the model, but in the decision layer, the exchange must deal with the model's output itself.
2.1 Exchange AI Features Viewed Through the Two Layers
The cases from Section 1 can be organized into execution and decision as follows. The way external models are connected to accounts is broadly similar, but how far access is opened and what is used to restrict it differs from one exchange to another.

In the execution layer, the point where exchanges actually compete is permission design. Binance ties API keys without withdrawal or transfer authority to a virtual subaccount separated from the main account. OKX keeps keys only on the user's device so the model cannot see them, and separately offers a read-only mode and a simulated trading mode. Coinbase has said it will have agents manage user capital only within an isolated portfolio, and plans to let users set spending limits and restrictions on trade size.
2.2 The More Execution Is Opened Up, the More Users Move Outside
Opening execution, that is, order functions, to a variety of tools cuts both ways for an exchange. Orders may increase as access routes multiply, but the exchange can no longer know what judgment lay behind those orders. The process of comparing markets and deciding what to buy and how much is completed entirely inside the external model, and only the conclusion reaches the exchange. In effect, the place where users meet the market shifts from the exchange app to the AI app.
This shift matters because the place where trust is built shifts along with it. Users build trust in a service while searching for information and checking the evidence. When that process happens elsewhere, the exchange only appears after the user has already made a decision, and the evaluation of the exchange narrows to whether execution was fast and how much the fees were. To the user, the exchange becomes a place that processes the results of decisions rather than a place that helps make them.
Of course, liquidity, execution quality, cost, and the reliability of asset custody still set exchanges apart, and if external models one day choose which exchange to send orders to, those will be the criteria then as well. But all of these criteria come into play after the user has made a decision. Opening execution alone does not give an exchange a way to build a relationship with users at the stage before the decision.
This is where the incentive for exchanges to provide the decision layer themselves comes from. The goal is not just to receive orders, but to give users a reason to stay within the exchange during the exploration that precedes an order.
A publishing-type decision layer differs from a conversational chatbot in its very form. Where a chatbot simply returned the model's answer to a user's question, the publishing type has a fixed format and update cycle, passes through evidence and verification set by the exchange, and goes out under the exchange's name. Bringing judgment into a place the exchange can control means bringing in the responsibility that comes with it. Which data to rely on, how much confidence to express, and how to filter out problematic results all become the exchange's job.
When connecting trading tools, the exchange's responsibility ends at permissions and the accuracy of execution. Even if a faulty request comes in, limiting in advance what it can reach is enough. When publishing analysis under its own name, the boundary cannot be drawn so neatly. What counts as a wrong explanation cannot be defined in a single line the way a permission can, and much of it only becomes apparent after the sentence has reached the user. And a wrong sentence leads to losses, regulatory risk, and damage to trust. So how much can be filtered out before it arrives? I will try to answer this by looking at traditional finance, which faced the problem first.
3. The Difficulty of Building a Decision Layer and Lessons from Traditional Finance
3.1 Why Building a Decision Layer Is Difficult
AI in traditional finance was adopted in stages: first for internal efficiency, then for compliance, and finally for customer-facing use. Crypto exchanges quickly followed the first two stages, and in customer-facing use they have already automated responses that do not involve judgment. At Coinbase's customer center, a Claude-based chatbot handles inquiries with compliance guardrails in place, and Upbit and Bithumb use AI to detect suspicious transactions and security threats. What remained empty was the third stage: information in the domain of judgment that the exchange sends directly to users under its own name.
This stage is hard to implement because the cost structure of failure is different. Errors in internal tools can be filtered by employees, and false positives from compliance AI can be corrected through human review. But flawed analysis delivered directly to retail users can affect their accounts without any such buffer. For that reason, judgment-type information must satisfy the following four conditions before it is provided.
- The boundary of advice: The moment generative AI says "buy this asset," it crosses into entirely different regulatory territory. Exchange AI has to walk a narrow path of providing information without becoming advice.
- Real-time requirements: Bank research moves on a daily cycle, but the crypto market moves around the clock, minute by minute. Analysis built on yesterday's data may already be wrong.
- The asymmetry of errors: Even an answer that is right 99 times out of 100 is hard to turn into a product in an environment where a single plausible falsehood leads to losses.
- Language: Since frontier models are trained mostly in English, maintaining the precision of financial language in Korean is harder than it seems.
Of these, the boundary of advice and the asymmetry of errors are conditions that traditional finance faced under regulation, while real-time requirements and language are additional burdens that crypto exchanges, and Korean exchanges in particular, must carry on top. Let us first look at how traditional finance handled the first two conditions.
3.2 How Traditional Finance Built Compliant LLM Services
In 2023, JPMorgan, Citi, Goldman Sachs, Deutsche Bank, and Wells Fargo restricted their employees' use of ChatGPT on compliance grounds. Three years later, JPMorgan has opened an in-house language model portal to more than 140,000 employees, and at Morgan Stanley, over 98% of wealth management advisor teams use an AI assistant.
This became possible not because the models grew more reliable, but because these firms built systems that did not require trusting the models. As of 2026, traditional finance has brought model deployment within compliance, and the following cases illustrate this.
- Morgan Stanley: Morgan Stanley built its own assistant agent on GPT-4. Rather than giving open-ended answers, however, the agent is restricted to answering only from roughly 100,000 research documents selected by the firm. Before and after deployment, the firm also repeatedly ran evaluations in which advisors themselves graded test questions drawn from their actual queries. AskResearchGPT, for institutional investors, follows the same structure and is designed to attach sources pointing back to the originals for answers that synthesize multiple reports. Debrief, a client meeting tool, summarizes meetings and drafts follow-up emails with client consent, but sending is decided only after the advisor reviews and edits the draft.
- JPMorgan: Through LLM Suite, JPMorgan provides external frontier models behind the bank's own control layer. More telling is IndexGPT, from May 2024. IndexGPT generates only keywords for investment themes, and the actual selection of securities was separated into a distinct procedure. If the model were to handle securities selection as well, it would enter the territory of advice.
- Goldman Sachs: In June 2025, Goldman Sachs rolled out the GS AI Assistant to all employees. The assistant is used for internal tasks such as document summarization, drafting, and data analysis, and was designed with different configurations for functions such as investment banking, development, research, and wealth management. Behind the assistant, models from OpenAI, Google, and Anthropic are connected alongside open-source models, allowing users to choose the model that fits their task. Goldman Sachs chose a design that lets users select the model but controls data access and permitted uses.
In short, what traditional finance built in order to deploy models within compliance comes down to the following four things.
- Selecting the material the model may reference
- Defining what the model will not do
- Evaluating outputs against expert standards
- Placing a final checkpoint before outputs are released
4. Building a Decision Layer: The Case of Bithumb and Surf

On September 2, 2026, Bithumb launched Bithumb AI, which it developed jointly with Surf. Surf is an AI partner to bring digital asset intelligence for financial institutions by analyzing on-chain and off-chain data in real time. In this collaboration, it took charge of the pipeline from data collection through verification and publication. This section first looks at the four features users see, then at how Bithumb AI was implemented to stay compliant, and finally at where this setup diverges from the analysis features of other exchanges.
4.1 The Four Features of Bithumb AI
Bithumb AI consists of four features, and the two asset-level features cover 450+ assets in the KRW market.
4.1.1 AI Signals


Every 5 minutes, it combines six technical and social indicators to present a short-term upward or downward signal and an overall AI score. The aim is to let users see the direction and changes of the trend on a single screen without having to compare multiple indicators and charts themselves. Bithumb stated that before adopting Surf, social indicators were excluded from backtesting because real-time market sentiment was difficult to reproduce. Since the collaboration with Surf, it said, market sentiment is reflected, and signal frequency and overall scores vary according to indicator weights that change over time.
4.1.2 AI Analysis

It generates an in-depth report for each asset by combining price, trading volume, on-chain data, and community reactions. Where Coin Signals condenses everything into a single score, the report spells it out along with the underlying evidence.
4.1.3 AI Daily Market Briefing


At 8 a.m. every day, it analyzes and summarizes the major market issues and also announces the day's economic calendar.
4.1.4 AI Suggested Questions

Every 30 minutes, it presents market issues of high user interest in a question-and-answer format. The service selects the questions itself so that users do not have to come up with them.
4.2 Surf's Four Design Choices for Keeping the Model in Check
When Surf published the design of this system on its engineering blog, it explained that the key was not a better prompt or a larger model, but owning the entire feedback loop around the model, in other words the "harness." Viewed against the four principles mentioned in Section 3.2, it becomes clear which principle each part of this harness is responsible for.
The first is selecting the material to reference. Every five minutes, Surf freezes market, on-chain, news, and social data into a single snapshot, runs it through more than 100 quality checks, distills it into a handful of explainable signals, and then hands it to the model. Letting the model search for evidence on its own slows processing and increases hallucinations. Where Morgan Stanley bounded the scope of answers with 100,000 research documents, here a bundle of data fixed at a point in time plays that role instead of documents chosen by people. The reason for fixing the point in time is a type of error peculiar to financial information. If a morning price movement and news released in the afternoon are placed in the same context, the later news can be described as if it caused the earlier movement. Checking only that each piece of data is factually true will not catch this error. That said, being placed in the same bundle does not mean the source data were observed at the same time, so collection time, the reference time of each source, and the delay of each data type must be managed separately.
The second is defining what the model must not do. Quantitative indicators such as price trends and trading volume are calculated in a predetermined way and assembled into scores, and the language model is only responsible for explaining those figures. After generation, an evaluation layer deterministically blocks outputs containing language that would amount to investment advice or mentions of competitors, and checks every figure against the source data. Through this, Surf aimed to keep the reports informative for trading without letting them drift in a direction that could be read as investment advice.
The third is evaluating outputs against expert standards. Bithumb AI's content is first generated in English and then translated into Korean. The system is designed so that an evaluation agent, trained on labels assigned by Korean-language reviewers at Bithumb and Surf, scores the translated sentences. This evaluation matters in financial writing because of the strength of conviction. "May rise," "is expected to rise," and "will rise" convey different information, and however smooth the translation, once a possibility turns into a forecast, the sentence no longer carries the same meaning. Surf states that Korean-language quality has reached a first-pass approval rate of 95% and a second-pass approval rate of 99%.
The last is the review process before outputs are released. When a quality problem is found in a generated piece, the error is spelled out in the next prompt and the piece is regenerated. If the model provider suffers an outage, the system switches to a model from another vendor and resumes from the point of interruption, and anything that still cannot be fixed is designed not to be released. Filtered-out failures are kept on record and shared with Bithumb's engineering team, which is also a point of difference from the systems of traditional finance. In the AI services of traditional financial institutions, the final checkpoint has consistently been handled by people. Manual review, however, is impractical for retail-facing market analysis that updates by the minute, so Bithumb and Surf moved that checkpoint to deterministic verification carried out in advance.
4.3 What Sets Bithumb AI Apart
What sets Bithumb AI apart can be divided into three levels.
First, users do not have to design their own questions. Conversational analysis tools commonly found at exchanges, such as TradeGPT in 2023, can be said to favor users who already know what to ask. Users who connect external agents directly likewise have to decide for themselves which data to call and what to ask. Bithumb AI shifts this burden to the provider through scores and reports in a fixed format and through suggested questions chosen by the service. It opted for a design that reduces not only the cost of generating analysis but also the burden on users of requesting and interpreting it.
Second, analysis and orders are integrated within a single app. As explained earlier, when analysis is delegated to external agents in the form of MCP, users encounter the market outside the app, and their contact with the exchange is kept to a minimum. Through the in-app experience, Bithumb AI deliberately creates reasons for users to make their decisions inside the app.
Third, everything from data selection to publication was designed as a single system. Bithumb is the first Korean exchange to regularly publish asset-level scores and reports, but news-summary analysis already exists in Robinhood's Cortex Digest and Upbit's AI news briefing. This means the type of analysis alone does not set Bithumb AI apart; what differs is the system behind it. Where a chatbot passed the model's output straight to the user, Surf states that its verification system, described above, filtered 300,000 failed outputs out of 2.5 million generation and verification attempts. This supports the view that such a filtering process is in fact necessary to build a decision-layer AI service at an exchange.
5. Exchange Intelligence Is Completed Outside the Model
Exchange AI started with chatbots inside the app, moved on to opening execution to the outside, and is now attempting to publish judgment under the exchange's name. Bithumb chose to build its execution layer with the AI Trade Kit and its decision layer with Bithumb AI, developed together with Surf. This article examined Bithumb's case in depth because Surf made its design public, which revealed in the most concrete terms the responsibilities and operational issues that come with building a decision layer.
What traditional finance showed first is that the condition for deployment was not trust in the model, but a system that did not require trusting it. That means selecting the material to reference, defining what the model will not do, evaluating outputs against expert standards, and placing a final checkpoint before anything goes out. The structure built by Bithumb and Surf carried these four over in the form of point-in-time snapshots, the separation of calculation from explanation, an evaluator trained on reviewer labels, and the withholding of publication.
There is no single answer to how far an exchange should go in making judgments. Bithumb chose to stay outside the boundary of advice. Quantitative indicators are calculated by rules, the model only explains, and any language that could be read as advice is removed before publication. Coinbase Advisor, by contrast, chose to step inside the boundary under the name of a registered adviser. The former manages responsibility by narrowing the scope of judgment, while the latter chose to bear the responsibility of advice within the regulatory framework. It is not yet clear which is more useful to users, but it is clear that both answers arose from the same question.
Having an analysis system of one's own ultimately means that the institution must be able to decide, and verify, which data to trust, how much confidence to express in its explanations, and when not to release a result. The same kind of work that goes into deciding what to allow for execution in the execution layer has to be done for each individual model output in the decision layer. The more information an exchange provides to support judgment, the more its competitiveness depends on the standards that govern those explanations rather than on their volume. What deserves attention in this attempt by Bithumb and Surf is precisely that it seeks to build those standards within the trading service itself.
The report is based on the independent research of the author sponsored/funded by Cybertino Inc. The author of this report may have personal holdings or financial interests in assets or tokens discussed herein. However, the author affirms that no transactions have conducted using material non-public information obtained in the course of research or drafting. This report is intended solely for general information purposes and does not constitute legal, business, investment, or tax advice. It should not be used as a basis for making any investment decisions or as guidance for accounting, legal, or tax matters. Any references to specific assets or securities are made for informational purposes only and should not be construed as an offer, solicitation, or recommendation to invest. The opinions expressed herein are those of the author and may not reflect the views of any affiliated institutions, organizations, or individuals. The opinions and analyses expressed herein are subject to change without prior notice. In addition, beyond the individual disclosures included in each report, Four Pillars, may hold existing or prospective investments in some of the assets or protocols discussed herein. Furthermore, FP Validated, a division of Four Pillars, may already be operating as a node in certain networks or protocols discussed herein or may do so in the future. Please see below links in the footer for FP Validated's participating network disclosures and for broader disclosure details.



