Combat skill scores 2.58 out of 5... Half of users saw it as a 'tool,' while only 18.5% recognized it as a 'teammate'

Built a Teammate, Got a Chat Partner: Krafton's First Scorecard for 'PUBG Ally'

Korean (한국어) Add us on

Krafton has published a technical report detailing the development process of 'PUBG Ally,' an AI companion for 'PUBG: BATTLEGROUNDS,' on the preprint repository arXiv. The 55-page report outlines the architecture and training methods, alongside user survey results from the live beta. In the survey, players rated Ally higher as a "partner who provides information and initiates conversation" than as the "teammate fighting alongside them" that Krafton originally envisioned.

Training data was collected over 28 days by renting out a local PC bang in South Korea. A total of 1,046 adults with over 10 hours of PUBG playtime participated for financial compensation, playing 38,956 sessions. Krafton swapped the on-device model six times during this period. The Net Promoter Score (NPS) rose from approximately +8%p in the early stage of collection to around +36%p in the final phase.

PUBG: 배틀그라운드 PUBG: BATTLEGROUNDS
©KRAFTON

"Built a teammate, but evaluated as a chat partner"

PUBG Ally is an AI character that plays matches paired up two-person squad style with a user. During the live beta, it was operated in the Arcade mode 'Ally Duo'. Ally converses with players via voice, navigates and engages in combat autonomously, and revives downed teammates. Krafton defined this as a 'Co-Playable Character (CPC)'.

The beta survey was conducted in 17 languages across 141 countries. The analysis was limited to respondents with confirmed gameplay records in Ally Duo. The Net Promoter Score, calculated by subtracting the percentage of negative responses ("would not recommend") from positive ones ("would recommend"), stood at +25.1%p, indicating a predominantly positive overall reception.

PUBG: 배틀그라운드 PUBG: BATTLEGROUNDS
©KRAFTON

Looking at individual categories reveals a different picture. All four gameplay evaluation categories failed to clear 3 points, the midpoint on a 5-point scale. Combat skill scored lowest among the four at 2.58 points, followed by situational judgment (2.88), response speed (2.93), and command execution (2.99). Conversation-related categories generally received higher scores than gameplay categories. In a question asking users to select up to two best aspects, "Information/Tactics" came first at 51%, while 28% selected "Conversation partner."

The same trend emerged in the question asking how users perceived Ally. Half of the respondents (50.0%) viewed it as a "tool." Respondents who viewed it as a companion—citing categories like "friend," "cute and wanted to take care of," or "special affection"—accounted for 31.5%, while recognition as a "teammate," Krafton's intended target, reached only 18.5%. In open-ended responses, key strengths cited included providing information such as enemy locations and directions, and making solo play feel less lonely. On the other hand, complaints included instances where Ally merely replied "Okay" without executing the action, or kept talking during late-game combat when players needed to focus on footsteps.

The report summarized that a positive recommendation intent coexisted with low evaluations of Ally as a combat teammate. Net Promoter Score was higher among respondents who accepted Ally as a teammate (+53.8%p) than among those who saw it as a companion (+46.7%p) or a tool (+40.2%p). While satisfaction was higher among users who felt Ally functioned as a true teammate, only about one in five users felt that way. However, survey respondents were players who voluntarily chose to play the mode, and Krafton did not disclose the absolute number of beta participants or survey respondents.

"$7 per match: On-device made it $0"

Ally runs its language model, speech-to-text (STT), and text-to-speech (TTS) entirely on the user's PC. The cost comparison table in the report reveals the rationale behind this decision. The median API call cost required for a single match was $5.14 to $7.10 depending on the language for Anthropic's 'Claude Opus 4.8,' and $3.70 to $5.17 for 'Claude Sonnet 4.6.' Google's 'Gemma 4 31B' cost $0.02 to $0.06.

PUBG: 배틀그라운드 PUBG: BATTLEGROUNDS
©KRAFTON

The report described Opus 4.8 as "impractical at the required scale." In a live game with high concurrent players, a structure where the company absorbs several dollars in inference cost per match is unsustainable. Because the on-device model runs on the user's graphics card, the API cost borne by the company is $0.

Speed was another reason. On an RTX 4060-class GPU, a single voice exchange took approximately 1.6 seconds for the on-device model compared to roughly 3.4 seconds for the cloud. The gap widens further in situations requiring multiple reasoning steps. When undergoing six rounds of language model inference, on-device finishes in under 5 seconds, whereas cloud exceeds 10 seconds. The recommended specification is 8GB of VRAM or higher. The 4-bit quantized language model occupies only 1.6 to 2.1GB to reserve system overhead for running the game.

Perceived quality among users also split based on speed. During the data collection period, a blind comparison was conducted without informing users which model they were playing with; the first on-device model (v1) was preferred over the cloud model (154 participants, 69.5% vs. 18.8%). The report drew a line, stating this result was "not evidence that a smaller model is superior, but rather the effect of response speed." At the time, cloud responses were slower than usual, and in a re-test with latency matched, the cloud model prevailed 52.4% vs. 31.1%. The takeaway is that in real-time games, a timely answer can dictate perceived quality more than a slightly more accurate one.

"From game AI teammate to robotics"

The report revealed that Ally's design is expanding into the robotics domain. Robot agent 'Ludi 0.1' from Ludo Robotics directly inherited Ally's core design principles and applied them to real-world conversation, memory, navigation, and object manipulation.

PUBG: 배틀그라운드 PUBG: BATTLEGROUNDS
©KRAFTON

The link bridging the two fields lies in their architecture. Ally was designed to separate the slow-thinking language model from a fast-reacting action control layer, allowing both layers to operate in sync. The language model interprets the user's speech and decides what to do. Actual movement, shooting, and reviving are handled by a separate layer that reassesses the situation with every game tick. The requirement to align speech with action while responding in real time applies identically to robots operating alongside humans.

Krafton presented the integration of a full-duplex voice model as a future objective. The goal is to move beyond the current push-to-talk system—where speech is only active while pressing a button—to allow users and the AI to interrupt each other mid-conversation. Starting in the constrained environment of Duo mode on the 'Sanhok' map, the AI companion leaves behind challenges regarding combat capability and user evaluation. At the same time, it serves as a stepping stone toward embodied agents operating outside of video games.

This article was originally written in Korean and translated with the help of AI. It was then edited by a native English-speaking editor. All AI-assisted translations are reviewed and refined by our newsroom. [Read Original]