|
|
A Recommendation Letter to Kong Studios' Project Team from a Game Development Novice (with Rich Gaming Experience)
— Also on How to Pair Multiple AI Models in the AI Era, and How to Use Codex to Push a Tower Defense Game from Concept to Art Production
To the Kong Studios Project Team,
Greetings. My name is Luo Yu. I am an ordinary person who knows nothing about the professional game development process, yet I have used AI to push a vertical-screen tower defense game, Buried Tides, all the way from a project proposal to mass art production. My engine, Godot, is installed on my D drive. I write code using Codex connected to DeepSeek (whom I affectionately call "Teacher D"), and I use gpt-image2 for image generation. This letter has two purposes: first, to organize my "what to do when" practical methods for reference by anyone in the team wanting to use AI to push projects forward; second, to systematically explain when and how to deploy the myriad of models—GPT, Claude, Gemini, Grok, DeepSeek, Kimi, GLM, Qwen, Hunyuan, MiniMax—and how to pair them most cost-effectively. I intimately understand the feeling of "tokens being too expensive to burn"—I also watched that video "Even Big Tech Can't Afford Tokens?" Microsoft and other giants are calculating AI costs, let alone us solo developers. Therefore, the underlying tone of this letter is not "which model is strongest," but "which model is most cost-effective at what time." Strong does not equal right; expensive does not equal should-use.
I. First, the Process: The Biggest Pain for a Novice Isn't Not Knowing How to Write, But Not Knowing Which Step They Are On
Originally, regarding "making games," I only had four words in my head: project initiation, prototype, vertical slice, expansion. Beyond that, everything was dark. What truly made things smooth was forcing myself (and forcing the AI) to cut the vague act of "making a game" into five pieces of bounded terrain: Stage 0 Project Proposal Documents, Stage 1 Greybox Prototype, Stage 2 Vertical Slice, Stage 3 Content Expansion, Stage 4 Polishing & Release. Each piece should only do what it's supposed to do; working across stages is almost certainly doing useless work. This is the main thread of this entire letter.
I ran this main thread completely in Buried Tides, and I can use real progress as a reference for the team.
In Stage 0 Project Initiation, the output is not code, but "decisions." Teacher D and I used one complete grand plan to chat out six volumes and over 40,000 words of specification documents: the main proposal, combat and numerical values, level design, UI layout contract, art style and asset list, and a "milestones and advancement suggestions" specifically written for novices. Values are defined once only in the numerical volume; other volumes only reference them, not rewrite them. The core of this stage is "settling all matters that need decision-making as early as possible," so that whoever writes the code later doesn't have to make any decisions for you.
In Stage 1 Greybox Prototype, the screen is all solid-colored blocks, not a single art asset is used, but the gameplay logic must genuinely run. Its sole mission is to answer "does this gameplay hold up, and is it fun?" I split the entire greybox into four sequential Goals: 1A Grid and Pathfinding, 1B Guardians and Gold, 1C Six Enemy Types and Ten-Level Spawn Table, 1D Chest Durability and Settlement. After each step, the AI was not allowed to just verbally say "done"; it had to actually run regression tests: numbers like the pathfinding steps shrinking from 27 tiles to 13 tiles when a giant rock is smashed—these were run by scripts, impossible to fake. During the greybox stage, I even stepped on a real bug—enemies walked right over the guardians. Teacher D didn't rush to move the pedestal; instead, it exhaustively enumerated the shortest paths under all 512 giant rock states, discovering that no grid cell could avoid the shortest path. The root cause was in the rule "deployment points can be passed through." So it changed the rule, rearranged the map, and wrote permanent regression tests to lock it down. This step made me thoroughly understand: regression tests are the only quality bargaining chip in the hands of a novice who doesn't understand code.
In Stage 2 Vertical Slice, it's finally time to "look good." I brought the first three levels to near-finished product quality: first generating 2x4 continuous sprite sheets for characters and enemies (protagonist casting, guardian attacking, normal mobs just walking), then generating full-screen combat concept art, environment asset sheets, and UI icon sheets. All assets uniformly follow the process of "magenta background generation -> cutout to transparent -> size normalization," using asset sheets to arrange multiple similar items at once to control style consistency.
Here is a blood-and-tears lesson I want to highlight with a blackboard eraser: Absolutely do not get itchy hands and start doing art during the greybox stage. The greybox is too ugly, and human instinct wants to beautify it, but the gameplay during the greybox stage will still change—maps will be rearranged, unit occupancy rules will change, cell sizes will go from 160 to 137 to 128, and art dimensions must follow. I reworked the concept art three times: the first version had one missing rock per barrier belt, the second version had all rocks but a hollow ground, and the third version was rigorous but the enemy snake direction was reversed. If I had rushed to produce art during the greybox stage, these reworks would have all been painted on finished assets, directly doubling token costs. Rules first, art second; small samples first, mass production second—this is the first principle of saving tokens.
II. The Two Modes of Codex: When to "Open a Plan" and When to "Open a Goal"
There are two advancement modes in Codex, and novices easily confuse them. Having used them extensively, I've summarized their respective boundaries.
Plan Mode is "only chat, don't do": the AI can read files, search, and run verifications that don't modify the repository, but it won't actually change any files. Its mission is to chat the specifications to "decision-complete" before acting—asking questions back until there is no need for on-the-spot decisions, finally spitting out a complete plan block, which you confirm and switch back to normal mode for execution. My criterion is one sentence: Only open a plan when the things to be moved span more than two volumes of documentation, or when entering a new stage. "Change the level 6 jellyfish from 6 to 8," "move the summon button up 20 pixels," "fix a type inference error"—such single-point changes can be said directly; opening a plan is forcing a process onto a simple task, purely burning tokens. But "elite monsters can also hit guardians," "change the battlefield from 6x8 to 7x10," "add a synthesis system"—these span volumes and change rules players are accustomed to, so they are worth opening a plan for.
Goal Mode is "setting a military order with an endpoint verifiable by the naked eye," spanning multiple turns and submissions. I only open a Goal when three conditions are met simultaneously: first, it takes multiple consecutive rounds to complete; if it can be wrapped up in one or two sentences, don't open it; second, there is an endpoint verifiable by the naked eye; acceptance cannot be "the AI says it's done," but must be a visual like "I press B to smash the giant rock and can personally see the cyan path shrink from 27 tiles to 13 tiles"—something screenshot-able or recordable; third, it can clearly state what is not being done this time,自带 boundaries, preventing the AI from diverging and adding drama during advancement, rolling the scope larger and larger—scope creep is the number one reason tokens are burned up and progress stalls in long AI tasks.
For a healthy Goal, I fixedly fill in four elements; the team can directly copy this skeleton: The first line is "Goal," one sentence; if you can't say it clearly, it means you haven't thought it through; the second line is "Deliverables," which must be countable—is it "3 scripts + 2 tests" or "8 sprite sheets"; the third line is "Acceptance Evidence," which must be "where I click, and then what I see." This is the key to turning "completion" from the AI's word into objective fact; the fourth line is "Explicitly Not Doing," which is the guardrail. My Goal to wrap up the greybox was written exactly like this: the goal is to make six enemy types run according to the behavior table and connect ten levels; deliverables are elite rock-breaking decisions, BOSS rock-throwing stun, ten-level spawn table, and corresponding regression tests; acceptance evidence is "in Level 5 you can see the murloc smash the giant rock and the whole field changes path, and in Level 10 you can see the BOSS stun the guardian for two seconds"; explicitly not doing chest settlement (that's the next Goal) and any art. A Goal given to AI is essentially a military order with acceptance criteria: whether it did it or not is not for you two to discuss; that acceptance evidence decides it.
III. Four "Must-Turn-Off" Iron Rules—Bought with My Tokens
At this point, I must list a few negative lessons separately, because they are the key to saving money, and the team should especially remember them when using Codex.
Before opening Plan Mode, go to Settings—Plugins—Skills and turn off brainstorming (requirements ideation skills). This type of skill is designed to repeatedly help you clarify goals, requirements, constraints, and solutions before development. Sounds good, but Codex's built-in Plan Mode is itself a mechanism for "chatting the specs to decision-complete while chatting." Two "repeatedly asking and clarifying" mechanisms stacked together will endlessly ask you back and divergently spread design options. Things that could be locked down in one step are flipped over and over, tokens gushing out, while progress stalls in the first round. The order is always: turn off brainstorming -> open plan -> chat out the plan -> switch to normal mode for execution. By the way, a sentence I often use with it: When you don't know what to do, open a plan and chat; chat until the direction emerges; but skills sometimes constrain the divergence of the plan, remember to turn them off when necessary.
Turn off prototype skills, turn off team (multi-agent parallel collaboration), and don't adopt the rank system/rank architecture. These are designed for "multi-person / multi-agent parallel" work. The core of the greybox prototype stage is highly serial and strongly dependent: if pathfinding isn't working, guardians can't be verified; if guardians aren't verified, enemy behavior can't be verified either. If you split this serial work to multiple agents in parallel, they will either be doing work whose prerequisites haven't been met, or duplicating labor, or waiting for each other—overhead doubles while progress doesn't increase. Similarly, the "architect/engineer/tester" rank division is organizational overhead completely unnecessary for the volume of one person plus one AI; applying it will only consume tokens on reporting and alignment rather than getting the game made. Rank systems, rank architectures, are not adopted.
Advance one window at a time. Since we're not doing parallel work and not doing ranks, the correct posture is: one window, focus on one Goal at a time (or even one change within a Goal), run it through, test it green, log it in CHANGELOG, then open the next. The biggest benefit of serial work is clean context—each window only contains the complete information of the current task, the AI isn't disturbed by the history of a bunch of parallel tasks, and you don't get lost either.
Converge these four rules into one sentence: In the initial stage of development, turn off all heavy mechanisms born for "multi-person parallel / repeated divergent clarification / organizational architecture," and return to the most plain way—one window, one Goal at a time, serial advancement, tests as the bottom line. Unless your tokens are especially, especially burn-resistant.
IV. General Principles of Model Usage: First Distinguish Two Dimensions
Next, how to pair models. This is the main body of this letter. I want to give the team a thinking framework first, then go through each series. Many people ask "which model is strongest" right away—this is the wrong way to ask. Using AI has two orthogonal dimensions: the model determines the capability ceiling, and the reasoning intensity determines the depth of thought. Using a lightweight model with maxed-out reasoning won't reach the effect of a flagship model; conversely, letting a flagship model handle "change a button color" is using a cannon to kill a mosquito, purely burning money. So the correct way to ask is: how high does the "ceiling" need to be for this task, and how much "depth of thought" is needed? Taking the GPT series in Codex as an example, the tiers are like this: small changes with clear goals use the lightweight tier with minimum reasoning; adding new pages, fixing common bugs that need a bit of planning use the balanced tier with medium reasoning (this is the default starting point for 80% of daily tasks); cross-file refactoring, security review, architecture design only go to the flagship with high reasoning; and "multi-agent parallel" Ultra tier is only for naturally splittable big tasks (like writing a microservice from scratch); non-splittable tasks like changing copy absolutely must not open Ultra. The two habits the team should most quit are: using the most expensive model for all tasks, and forcefully using the cheapest model to tough out complex projects.
V. GPT Series: How to Divide the Three Tiers of Sol, Terra, and Luna
OpenAI changed naming from version numbers to celestial names this generation. In Codex, it's mainly the three GPT-5.6 brothers. I'll match them to uses.
Sol is the flagship, strongest capability, specifically reserved for large-scale refactoring, difficult bug root cause investigation, cross-file transformation, architecture design, high-risk code review, and long-duration autonomous agent tasks. API is about $5 input, $30 output per million tokens (doubling for long contexts over 272k), expensive for a reason, don't use it for daily work. Terra is the balanced tier, and the one you should use every day—daily business feature development, interface modifications, regular bug fixes, adding unit tests, writing documentation, medium-complexity tasks all go to it; after its 20% price cut in late July, input $2, output $12, balancing capability, speed, and cost, even usable in Codex free tier and entry subscriptions. When you ask "when to use Terra," the answer is: except for the heavy tasks of Sol above and the light tasks of Luna below, the vast middle ground all goes to Terra. Luna is the lightweight tier, fastest and cheapest, after an 80% price cut only $0.2 input, $1.2 output, specifically for repetitive labor like variable renaming, style adjustments, code formatting, adding comments, batch extracting information, quick Q&A, but it doesn't support Ultra.
One-sentence quick reference: Change copy, change colors, format—find Luna; add pages, fix common bugs, write tests—find Terra; refactor large functions, cross-file debugging, security review, difficult root causes—go to Sol. When is daily? Terra is daily. When to find bugs and fix them? Common bugs use Terra, recurring intractable ones escalate to Sol. When to reason and write code? Writing complex logic, architectural code requiring trade-offs uses Sol with high reasoning. When to diverge thinking? Divergence itself should use Plan Mode to "chat," not rely on a model to think hard—the model only provides material, divergence comes from the process of you chatting the plan out. In the previous generation, GPT-5.4 was good for reading large projects, gnawing on legacy systems, 5.4-mini was one-seventh its price and could handle some of Luna's work; if you can still select them in your account, they are also money-saving supplements.
VI. Claude Series: First Tier for Writing Code and Long-Horizon Tasks
Anthropic's current generation focuses on programming and long-duration agent tasks. The main force is Opus 5.5, with a more expensive Fable 5.1 above as the capability ceiling, and Sonnet 5 below for cost-effectiveness. Opus 5.5 was priced at $4 input, $20 output per million tokens at release, about 40% cheaper than the previous generation, either leading on multiple real programming benchmarks or matching GPT-5.6 Sol at about one-third the cost—its relationship with the GPT series is not who replaces whom, but mutual quality inspection: let Claude write, let GPT review, or vice versa. Cross-validation is much more stable than trusting one alone. When to use Claude? Complex multi-file changes, long-horizon tasks where the model needs to work continuously for hours without getting distracted, and "when you want another model's judgment on the same code." Daily simple tasks are too expensive for it, not cost-effective, leave them to the cheaper options below.
VII. Gemini and Grok: One Relies on Context, the Other on the Internet
Gemini's current generation (Gemini 3.1 Pro hit a high score of 80% on programming benchmarks like SWE-bench, even higher than the Claude series) has super-long context as its biggest ace. When to use it? When you need to feed an entire repository, hundreds of thousands of lines of code, or a full set of documents in one breath for it to read through and then answer, its long-context advantage is most obvious; daily multimodal, reading images and videos are also its strengths. Grok is xAI's, strong at real-time internet access and "daring to speak, divergent, with personality," suitable for opening when you need to grab the latest information, do viewpoint collision, or brainstorm; but its price is not cheap, Grok 4.6 is about $3.4 input, $40 output, doubling for long context, so it should appear in narrow scenarios of "need internet + need divergence," not as a daily main force.
VIII. DeepSeek, Kimi, GLM, Qwen, Hunyuan, MiniMax: The Domestic Camp's Cost-Effectiveness Hinterland
The highlight is here—if the Kong team wants to control costs, the answer is here. My personal judgment is very clear: using DeepSeek daily is absolutely cost-effective; this is not sentiment, it's forced by the price list.
DeepSeek now has V4-Flash and V4-Pro tiers, 1 million context, supports thinking/non-thinking switching, and directly provides an Anthropic-format API endpoint—meaning it can be connected into Codex like Claude as "Teacher D" (I've been using it this way throughout). Price: V4-Flash miss input 1 yuan, output 4 yuan per million tokens, cache hit input as low as 0.02 yuan; even the strongest V4-Pro at peak is only 9 yuan input, 27 yuan output. Compared to foreign flagships costing tens of dollars per million tokens, it's orders of magnitude cheaper. And it has peak/off-peak pricing: weekday daytime is peak, evenings and weekends are off-peak, directly half price. So my money-saving habit is—schedule heavy work to off-peak hours (night, weekends) to run; the same money does double the work. When to use DeepSeek? Daily development, writing code, fixing bugs, reading documentation, running long Goals, use it most of the time, cost-effectiveness is unmatched. When to find and fix bugs? DeepSeek is completely sufficient for common bugs, and cheap enough that you can confidently let it try multiple rounds.
GLM (Zhipu) current generation GLM-5.3 is about 8 yuan input, 28 yuan output, but its Flash tier GLM-5.3-Flash is as low as 0.8 yuan input, 2.8 yuan output, same order of magnitude as DeepSeek-Flash, a swappable cheap main force; Zhipu this generation also sells "agents optimizing infrastructure in reverse," long-horizon agent capability is progressing, worth being DeepSeek's second backup, cross-validation. Kimi (Moonshot AI) K3 has strong capability, long text processing is its signature, but note its output is very expensive (output about 100 yuan per million tokens), so Kimi shouldn't be used for "chatty" long-output tasks, more suitable for "read a lot of long material, give a concise conclusion" scenarios with much input and little output. Qwen (Alibaba Tongyi) is also a Max flagship + Flash lightweight dual tier, Qwen3.8-Flash input 0.8 yuan, output 2.7 yuan, also a cheap tier, Chinese and long context are traditional strengths, suitable for massive Chinese text processing and daily Q&A. Hunyuan (Tencent) Hy4 preview input 6 yuan, output 18 yuan, between cheap tier and flagship, convenient integration in Tencent ecosystem. MiniMax has a group of followers in voice, anthropomorphic dialogue, and role-play language sense; if your project has NPC dialogue, voice, emotion-oriented copy needs, it's a characteristic supplement, but pure programming is not its home turf. StepFun's Step 5 and other new open-source flagships are also pushing "unit intelligence cost" down, claiming single-task cost is only one-eighth of Opus and planning to open source; the team can put it on the watchlist for technology selection.
IX. A Concrete Pairing Plan for the Kong Team
Converge the above into four default rules the team can execute tomorrow. First, default main force uses DeepSeek (or equivalent GLM-Flash, Qwen-Flash); daily development, coding, bug fixing, long Goals all go to it, heavy work scheduled to off-peak hours as much as possible. Second, only when the cheap model fails two or three rounds in a row, temporarily upgrade to GPT-Sol or Claude-Opus for the tough battle, immediately downgrade after the battle—treat flagships as "expert consultation," not "daily clinic." Third, use a second model as quality inspection: important code let DeepSeek write, GPT review, or vice versa, cross-check once before merging, much more stable than trusting one model, and only spending a small amount of flagship tokens. Fourth, Gemini specializes in super-large context reading, Grok specializes in narrow scenarios needing internet and divergence, Kimi specializes in long material read-much-write-little, MiniMax specializes in voice and anthropomorphic dialogue—let them each play to their strengths, no one goes on daily. The essence of this combination is: spend the most expensive compute precisely on the few things "only flagships can do," and everything else goes through cheap models and off-peak hours.
X. One More Thing More Token-Saving Than "Choosing a Model"
Finally, a saving method I stepped on that's easily overlooked: For image generation and critical reasoning, don't grind hard in a relay station, high-latency, suspected degraded-intelligence environment. I initially used a relay station's model to run gpt-image2 image generation, repeatedly redoing it amid high latency and obvious quality degradation, extremely token-consuming and unable to produce stable results. Image generation wants stable picture quality and style consistency, which is exactly sensitive to the model's "full-bloodedness"—a makeshift line will make the same prompt repeatedly fail, and you think it's a prompt problem, but actually the environment is dragging you down. It's fine if it works, but don't grind on an obviously degraded line; when it's time to use a stable direct connection environment, don't save that little price difference. The essence of saving tokens is not always picking the cheapest, but not wasting tokens on rework—and the biggest source of rework is acting before rules are set, and grinding hard in a degraded environment.
At the end, allow me to say something unrelated to technology. Last night I dreamed of the world in '23 and '24, fighting the invasion leader, a space war, reaching out to hard-material battleships, partners guarding the pass, a forest to rush to before the deadline. I woke up and smiled—probably during the day I was designing tower defense "guard the chest, endure wave after wave," and at night my brain fought a "guard the world" battle for me. Making games and dreaming are actually the same thing: you give it a set of rules, a deadline, something that must be guarded, and it grows levels, monsters, and countdowns on its own. I hope this letter gives the Kong team not just how to save tokens, but this sense of certainty in breaking the vague "I want to do something" into steps that are judgeable, verifiable, and iterable. May we all look forward to the new world, and truly, version by version, make it happen.
The above suggestions are for the team's reference. Specific prices are subject to each manufacturer's real-time official website (my numbers are as of September 2026; models and prices change every month, please refer to the latest).
Looking forward to the new world.
Note: If using the DeepSeek model, the request body limit for a single task window is (plain text about 880kb), otherwise a 400 error.
Inline images (sending images in the chat box, base64), limit is 48mib, otherwise a 413 request body limit error.
File reference images limit is 200mib, otherwise 413.
Currently, it seems that once the request body reaches the limit, you can only switch to a new window...
Sincerely,
Luo Yu
September 23, 2026
thank you~
첫댓글 ♡
안녕하세요. 洛雨 기사님
AI를 활용한 개발 경험과 파이프라인 구축에 관한
소중한 의견을 말씀해주셔서 감사합니다.
보내주신 제안과 유용한 인사이트는 모두 관련 담당 부서로 전달하여
꼼꼼히 검토할 수 있도록 하겠습니다.
감사합니다.
I forgot to mention, if you plan to use the inexpensive DeepSeek daily, I recommend setting a 400k context limit for long inferences and a 128k context limit for short inferences. That's about it. Thanks.