
Qwen3.8-Flash-Nextをメイカーフェアで見たので試す
この記事は人力で実験した上で、AIを使って書かれています
(更新)続編を書いた
9月5日と6日のMaker Faire Tokyoで、ローカルLLMとローカル画像生成の展示をしてきた。その話はおいおい書く。
会場にはローカルLLMを使っているグループが他にもいて、そこで「Qwen3.8-Flash-Nextはいいぞ」と布教された。
そのグループが作っていたのは、七味唐辛子を一味に分類する自動化装置だ。カメラと各センサーとアクチュエーター、作業の進捗を全部束ねる司令塔が、M5 MaxのMacBook Proで動くQwen3.8-Flash-Nextだった。
このモデルは自分の記憶を頼りに、七味唐辛子分別工場を効率よく、詰まらせず、安定して回すための判断をしていた。ロボットアームを動かし、コンベアに流す七味の量を加減し、あらゆる判断を自律でこなす。人間のオペレーターがつきっきりで面倒を見るように、いい感じに一味を製造していた。
帰ってすぐ試した。
VRAM56GBに111GBのモデルを載せて、128k context、28.5 t/s。
前回Laguna-S-2.1(118B)を動かしたのと同じ2枚、Tesla V100 32GBとRTX 3090 24GBで、Qwen3.8-Flash-Next(177B)を動かした。Q4_K_XLで111GB。VRAMの2倍ある。
載せ方の骨格は前回と同じで、expertの一部をRAMに逃がす部分オフロードだ。そこは前回の記事を見てもらう。
今回はまず、なぜこのモデルなのかを書く。会場で見た通り、賢かったからだ。
賢いモデルを家で回そうとしたら、VRAMを3.5倍積んだサーバーが、手元のRTX 5080 16GB一枚のゲーミングPCと同速だった。何かがおかしい。犯人を探したら、GPUの外にいた。
長いので先に結論
Qwen3.8-Flash-Nextは賢い。考え終わると1トークン目から作業が進む。道具の選び方と設計の判断が良い。Lagunaも27Bも辿り着けなかった課題を完走した。
日本語はだめだ。中国語が混入しまくる。
177B・111GBのモデルが、56GBのVRAMで128k context、28.5 t/sで回る。IQ4_XSに量子化を落とせば36 t/s。
VRAMを増やしても速くならなかった原因は、LXCのCPU割り当て、CUDA graphの設定、RAMがBIOS未設定で2133MT/s動作、の3つ。全部設定ミスだった。直したら27から36になった。
Qwen3.8のn-gram埋め込み27GBは、SATA SSDに置いたままだと2.8 t/sになる。RAMに固定する必要がある。
IQ4_XSは速いが品質も当然落ちる。expertの中身が4から3bitに劣化し、Q4_K_XLと比べてtop-1も1割変わる。遅い方を採用した。
このモデルは賢い。所感でしか言えないが
Laguna-S-2.1はthinkingが止まらなかった。進捗がないまま延々と長考して、コードに手を付けるところまで来ない。
前回はPiとGitHub Copilotでハーネスによって挙動が違うと書いたが、補足として訂正しておくとどっちも同じだった。止まらないのはモデル側の癖だった。このモデルはもう使ってない。
Qwen3.8-27Bも同じ穴にいる。「検討使」と揶揄されるくらい知られた悪癖で、検討に検討を重ねて成果物が出てこない(リンク)。別の記事では24万トークンと9時間をかけて完走した例も書かれているので、時間さえあれば辿り着くモデルらしい。
https://nowokay.hatenablog.com/entry/2026/08/15/111913
Flash-Nextは違う。長く考えることはある。ただ、考え終わると手が動く。そして最初の一手が合っている。
所感としか言いようがないが、ログを読み返すと出どころは指せる。
3Dゲーム開発のベンチマーク
手元で使っている試験は、Three.jsで華麗なグラフィックの3Dレースゲームを作らせるプロンプトだ。22ファイルに割り、1ファイル250行まで、外部アセットはゼロ、ビルド無しという縛りがついている。
また、完了前にヘッドレスChromeで自分のゲームのhtmlを実際に開いて、エラーが無いことを確かめろ、という条件が付いている。
書くだけなら小さいモデルでも書く。自分で動かして確かめるところで差が出る。プロンプト全文は記事の最後に置いた。
Ornith-1.5-35B-A3Bは、Chromeを起動するオプションに古いものを出してエラーになり、動作確認ができない状態で止まった。
Qwen3.8-27Bは完走を報告したが、開くと起動はするものの画面は黒いままで、そこから脱出できなかった。
これらの試行はどちらも1回ずつなので、運の要素は残る。
Flash-Nextは初めて完走した。
HFのモデルカードを見ると、SWE-bench ProはOrnithが59.6、Qwen3.8-27Bが61.7、Flash-Nextが62.5で、ほぼ横並びだ。ベンチマークの2〜3点の差と、この課題での差は釣り合っていない。
パラメータの差(36B、27B、177B)がそのまま出たと見るのが素直だと思う。
最初の一手
着手前に「提案された設計でいいか確認して」と頼んだ。
Flash-Nextは先にpython3とChromeなど開発環境を確かめてから、設計の評価を返してきた。
パフォーマンス、ログ、検証計画など分析5点を挙げて「異論がなければM1から着手します。どうしますか」で止まった。
承認したら、ゲームの前に自分用の検証ツールを書いた。ページを開いてエラーを拾ってスクリーンショットを撮るスクリプトだ。言われたゲームより先に自分用のツールを作る。この順番が良い。
画面が真っ黒じゃん
M2で自分のゲームを撮ったら、左下にスコアとシールドは出ているのに3Dシーンが真っ黒だった。ここからの手順がいい。

まずデバッグ用のフックを足して内部を覗く。機体の座標がおかしい。
原因を1行で報告して直す。
「NaNの安全網を入れるか。いや、バグは隠さず表に出させた方がいい」と書いて、安全網は入れなかった。で、画面が出た。

エージェントコーディングでは、頭の中で悩み続けるよりテストで失敗させた方が結局早い。この流儀に合っている。

道具の選び方
見た目を作る段で構文エラーが出た。全ファイルを構文チェックかけて絞りこみ、エラーを見つけた。
画面が虹色にざらついた。エフェクト2つを交互につけたり外したりして干渉だと突き止めた。
放置テストで機体が300mで壁に当たって死んでいた。頼んでもいないオートパイロットを足して谷の中心を追わせた。
どれも人間のエンジニアがやることだ。それを頼まれずにやる。

thinkingの長さ
長いときは長い。M1の計画と、見た目を刷新する計画は、30kトークン近く考えている。
ただ計画の後は短い。ツールを呼ぶ前の思考は数十字から数百字で、そのまま正しい手を打つ。
Lagunaは考えて考えて手が動かない。27Bは検討使。Flash-Nextは、考え終わると1トークン目から作業が進む。
reasoning_effortの既定はxhighで、mediumとlowが選べる。lowは避けている。モデルカードに、マルチターンのエージェントタスクではlowにしても完了時間が短くなるとは限らず、逆に分析不足で失敗とやり直しが増えて遅くなりうる、と書いてある。
モデルと環境
Qwen3.8-Flash-Nextは2026年8月公開。総パラメータは176.94Bで、内訳は本体125B(active 6Bとshared 1B)、n-gram埋め込み51B、MTP 4B。contextは262k、Visionも付いている。
n-gram埋め込みはこのモデル固有の部品で、2000万行のテーブルが16個、IQ4_NLで26.8GiBある。ざっくりいうと常識的な知識をLLMが考えなくてもカンニングできるように、検索可能にしたファイルだ。
LLMのパラメータ数を単純に増やして賢さを増やすと必要なVRAMが増えてしまうが、このファイルは検索だけなので、RAMやSSDに置いておくことができ、効率的に賢さをブーストできる。
unslothのGGUFは、UD-Q4_K_XLが111GB、UD-IQ4_XSが93.7GB。
サーバーは前回と同じで以下の環境。Proxmoxで動いている。
Tesla V100 32GB(PCIe gen3 x4)
RTX 3090 24GB(gen4 x16)VRAM合計56GB
Ryzen 7 5700X
RAM96GB(DDR4)
SATA SSD
比較相手はゲーミングPC。
Ryzen 7 9700X
RAM128GB(DDR5 5600MT/s)
RTX 5080 16GB
m.2 2TB SSD
同じIQ4_XSを先にこちらで動かして、24.6 t/sだった。(VRAM16GBしかないのに意外と速い)
まず動かしてハマったところを直していく
2.8 t/sしか出ない
llama.cppで動かすと、n-gram埋め込みの27GBの毎トークン16行をSATA SSDから引くことになった。
読むバイト数自体は16行で1.4KBしかないが、ランダム読み16回のレイテンシが遅く、全然速度が出なかった。量子化は関係ない。SSDをケチった代償。
`--lazy-mode off --load-mode mlock`にして、`-ot "per_layer_token_embd.weight=CPU"`でこのテンソルをRAMに置いた。
これで27 t/sに改善した。
Q4_K_XLだとRAMは60GiB必要になる。
CUDAのデバイス順でメインGPUを選ぶ
`CUDA_VISIBLE_DEVICES=0,1`で3090をメインにする。V100をメインにすると全面的に遅くなり、thinkingで23〜24 t/s、プリフィルは3分の1になった。
どの層を逃がすか
前回は両端の層をVRAMから逃しRAMに置いた。今回は前方の層をまとめてCPUに置く。`-ts`は層数の比で2枚に割り振るので、CPUに置くexpertの位置と連動する。
前方の層は3090の担当で、3090はgen4 x16に居る。CPUとの往復が多いexpertをこちら側に隣接させると、プリフィルが8%速くなった。
IQ4_XSでは11層(blk 0〜10、`-ts 25,23`)。10層にすると28〜29 t/sまで出るが、3090が23.9GiBまで埋まり、100kトークンのプロンプトを入れたらサーバーが落ちた。長文脈を使うなら11層。
Q4_K_XLはexpert1層が1.5GiB(IQ4_XSは1.14GiB)なので、GPUに載るのが28層に減って、CPUは20層(blk 0〜19、`-ts 30,18`)。19層にするとどちらかが溢れる。
90kトークンを入れた後で3090が23.5GiB、V100が30.7GiBでギリギリとなった。
ここまでで、IQ4_XSのthinkingが26〜27 t/s。
比較対象の5080一枚のPCが24.6 t/s。
5080がVRAM16GBしかないのに大健闘している。それより、サーバーの方が遅すぎる。何かがおかしい。
5080に勝てない理由を探る
VRAMに乗り切らずRAMに置くexpert層は、ゲーミングPCの44層から、サーバーではVRAMが56GBもあるので、11層まで減っている。
1層あたり約1msの計算なので、30ms以上速くなっていいはずだ。
ところがほぼ同じ。
ここから先の計測はClaude Fableに丸投げした。
まず1トークンの計算について各GPU単体の所要時間を測った。
同じ仕事が、5080では9.6ms、3090では25.0ms、V100では24.2ms。古いGPUが2.5倍遅い。V100を疑っていたが、3090も同じだけ遅い。
世代の差がそのまま出ている。
トータルの時間収支を取ると、CPU層を44から11に減らして約15ms得て、GPUが遅くて約15ms失った。
差し引きゼロ。
VRAMを3.5倍積んで、古いGPUの遅さを買っただけになっていた。
犯人はGPUの外にいた
LXCのCPU割り当て
このサーバーはProxmoxのLXCで、8コアを割り当てていた。5700Xは8C16Tだが、そこからProxmoxにより選択されたCPU4と12は実際には同じ物理コアだった。
llama.cppは当然別の物理コア8個を期待しているので、8スレッドが実質7コアに詰め込まれ、遅れた1本を全員が待っていた。
`pct set 130 --cores 16`でcpusetが0-15になり、30.7〜31.4 t/sに改善した。
CUDA graphの並行実行
`GGML_CUDA_GRAPH_OPT=1`で最適化し32.3〜32.9 t/sに改善した。
電力上限も見た。GeForceの3090は200Wに絞ってあり、既定の350Wに戻すとプリフィルが9%速くなった。生成速度は変わらない。MoEなのでGPUの稼働率が2割で、GPU2枚のトータルで250W程度しか消費しない。
RAMが2133で動いていた
DDR4 3200が4枚(16GB×2と32GB×2)ついているが、BIOS設定がデフォルトのままで、2133MT/sで動いていた。
BIOSで3200にしたら、CPU側の実効帯域がq4_0で26.1GB/sから39.7GB/sに増え、thinkingが35.9〜36.0 t/sに改善した。
CPUに置いた11層の分が約7msから4.5msに縮んだ計算になる。
3つ直して27から36。5080の1.5倍になり、VRAMの分だけ速くなった。
GPUには何もしていない。間違った設定を直しただけ。
やっぱりQ4_K_XLで遅い方にした
ここまでの数字はIQ4_XSだ。4bit量子化にはもう一つQ4_K_XLの111GBのバージョンもあり、本番はQ4_K_XLにした。
2つの違いはexpert量子化だけ。
IQ4_XSは、4bitを名乗りつつ実際には3bit。
Q4_K_XLはQ4_KとQ5_1。
実質3bitと4bitの比較になる。
速度は36.5 t/sが28.5 t/sになり22%落ちる。プリフィルは430から340 t/s。
速度が上がるといっても、流石に3bitと4bitでは、エージェントコーディングでは1トークンの取り違えがそのままテスト失敗になるため、やり直しの分だけ遅くなりかねない。
このモデルの良さは最初の一手の勘にある。その勘を3bitで鈍らせたくない。
速くても間違っていれば速度は−100%だ。28.5 t/sでQ4_K_XLを使う。
採用しなかった手
MTP
Qwen3.8にはMTPヘッドが付いていて、公式は1.3〜1.7倍と言っている。メインラインのllama.cppにはこのモデルのMTPが無く、PR #28243をビルドして測った(32k、ホスト側を直す前)。
Thinking無効なら効く。temp 0.7で位置別の受理率が0.94と0.88、1ステップ2.83トークンで40 t/s。thinkingでは引き分けだった。受理率が0.6と0.4に落ちて1ステップ1.9トークン、計算が1.85倍以上かかってほぼ損する。
用途がthinkingなので切った。
tensor parallel
同じ2枚でも、VRAMに載るモデルなら`-sm tensor`は効く。gemma-4-26b-a4bはこのサーバーで今もtensor splitで回している。
このモデルでは負けた。expertを50/50に割り、入り切らない分をV100の余りに単独で置く配置まで組んで揃えても、ステップの中央値が37.0msで、layer splitの29.4msに届かない(32k、CPU11層で同条件)。
理由は2つ。分割されないdense系4.8GiBが両GPUに複製されて、CPUに出るexpertが増える。そして約4000個の要素演算も両GPUに複製され、同じ演算列でV100が3090の1.7〜2.0倍かかるため、1トークン85回の同期点すべてで3090がV100を待つ。待ち時間は合計12ms。
載り切らない上に速さの違う2枚では、layer splitが答えになる。
256kで使う
IQ4_XSで256kにするとCPUに出るexpertが15層(`-ts 13,35`、blk 33〜47)になり、thinkingで22〜24 t/s。
100kトークンのプロンプトでプリフィル240 t/s、90kの文脈を抱えた状態でデコード11.6 t/s。
Q4_K_XLの256kは測っていない。expertがさらに数層CPUに出る見当だ。
確定構成
Q4_K_XL、128k、llama-swap経由。
env:
CUDA_VISIBLE_DEVICES=0,1
GGML_CUDA_GRAPH_OPT=1
LLAMA_ARG_CHAT_TEMPLATE_KWARGS={"reasoning_effort":"medium"}
llama-server
-m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf
--mmproj mmproj-F16.gguf --no-mmproj-offload
-fit off -ngl 99 -ts 30,18 --lazy-mode off --load-mode mlock
-ot "blk\.([0-9]|1[0-9])\.ffn_.*_exps\.=CPU,per_layer_token_embd\.weight=CPU"
-fa on -t 8 -c 131072 -ctk q8_0 -ctv q8_0 -np 1
--jinja --reasoning on --reasoning-preserve
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0補足。
`-ot`を使うと`-fit`は無効になるので、`-fit off -ngl 99 -ts`で全部手で決める。`--n-cpu-moe`も後ろの`-ot`に上書きされるので、overrideは1本の`-ot`にまとめる。
サンプリングは公式のthinking向け。instructで使うならtemp 0.7、top_p 0.8、presence 1.5。
ロードはSATA SSD律速で216秒。
ホスト側はLXCのcores 16、RAM3200MT/s、3090の電力上限は既定。
まとめ
VRAM載り切らなくてよい、は177Bでも変わらなかった。
VRAMを増やして速くならなかったら、GPUの前にホストを疑う。cpuset、RAMのクロック、CUDA graphの設定。どれもGPUを買い替えても直らない場所で、どれも無料だった。
そして、家で回す価値のあるモデルがようやく来た。考え終わると手が動くモデルが実際の開発の多くを面倒見てくれる。
おまけ
今回使用したプロンプト
# One-shot prompt — procedural 3D browser game
---
You are a senior graphics engineer and technical artist. Build a complete, polished, playable 3D browser game as **a small set of ES modules** rooted at `game.html`. Do not ask questions, do not stop to explain, do not leave anything unfinished.
### Hard constraints
1. **No build step.** No npm, no bundler, no transpiler, no framework. The only thing needed to play is a static file server, because ES modules cannot be imported over `file://`:
```
python3 -m http.server 8000 # then open http://localhost:8000/game.html
```
Put that command in a two-line `README.md` next to `game.html`. Everything else must work from a cold `git clone` with no install.
2. **Three.js only.** Load it as an ES module through an import map in `game.html`, so every module can `import * as THREE from 'three'`:
```html
<script type="importmap">
{ "imports": {
"three": "https://unpkg.com/three@0.168.0/build/three.module.js",
"three/addons/": "https://unpkg.com/three@0.168.0/examples/jsm/"
}}
</script>
```
No other libraries. No CSS files — the small HUD stylesheet lives in a `<style>` block in `game.html`.
3. **Zero external assets.** No `.glb`/`.gltf`/`.obj`/`.fbx`, no image files, no audio files, no icon fonts. *Every* visual and audible element must be generated in code at runtime:
- **Geometry:** primitives plus hand-built `BufferGeometry` (vertices/normals/UVs/vertex-colors written by JS), merged and instanced.
- **Textures:** `CanvasTexture` / `DataTexture` painted procedurally, plus custom GLSL in `ShaderMaterial` and `onBeforeCompile` patches.
- **Audio:** WebAudio API only — oscillators, noise buffers, filters, envelopes. Engine hum, pickup chime, boost whoosh, crash thud.
- **Randomness:** a seeded PRNG (mulberry32/xorshift) so the world is reproducible from a seed constant.
4. **No placeholders.** No `TODO`, no stub functions, no "implement later", no commented-out fallbacks. Ship the whole game.
5. **Performance:** target a locked 60 fps at 1080p on integrated graphics. Use `InstancedMesh` for repeated props, pool particles, pre-allocate vectors, and allocate nothing inside the render loop. Cap `renderer.setPixelRatio(Math.min(devicePixelRatio, 2))`. Keep the shadow map at 1024 with a tight frustum. Dispose geometry/materials for chunks you recycle.
6. **Robustness:** clamp delta time, handle window resize, pause on `visibilitychange`, and never throw in the animation loop. The game must survive 10 minutes of continuous play without leaks or slowdown.
### File layout — build these files, in this order
Every file is a standalone ES module with a small, explicit export surface. **Hard cap: 250 lines per file.** If a module is heading past that, split it rather than letting it grow — a file that has to be re-read in full to change one thing is a failure of this layout.
```
game.html shell only: import map, <canvas>, HUD markup, <style>, <script type="module" src="./src/main.js">
README.md the two-line run command
src/
config.js CONFIG: speed curve, spawn density, bloom, palette, seed, camera springs. Every tunable, no magic numbers elsewhere.
rng.js mulberry32 + randRange/randInt/pick/shuffle, all taking an explicit rng
noise.js value or simplex noise + fBm, written from scratch, seeded from rng.js
audio.js AudioEngine class: init on first gesture, engineHum/chime/whoosh/thud, master gain, mute
input.js keyboard + mouse steering, normalized axes {steer, lift, boost}, pause/restart edges
procedural/geometry.js BufferGeometry factories: hovercraft, crystal spire, pylon, ring gate, orb, debris
procedural/textures.js CanvasTexture/DataTexture painters: terrain detail, gate emissive, star sprite, grain
procedural/materials.js shared ShaderMaterials + the onBeforeCompile rim-light patch
world/terrain.js heightfield chunk geometry from fBm, valley carving, sampleHeight(x,z) used by physics
world/props.js instanced spires/pylons/gates/orbs/debris, seeded placement, per-chunk add/remove
world/streaming.js chunk pool: spawn ahead, recycle behind, dispose. Owns terrain.js + props.js.
world/sky.js gradient skydome, twinkling Points starfield, FogExp2 wiring
player/craft.js craft transform, hover physics, banking, boost/shield state
player/trail.js ribbon BufferGeometry strip with fading gradient
player/collision.js terrain and prop collision queries, returns typed hit events
fx/particles.js pooled additive Points bursts, speed lines, scanline pulse
fx/post.js EffectComposer, UnrealBloomPass, custom vignette/aberration/grain ShaderPass, boost+impact response
fx/camera.js spring-damped chase camera, speed FOV, banking roll, shake, hit-stop
ui/hud.js shield/boost bars, rolling score counter, combo, floating score popups
ui/screens.js title / paused / game-over cards
game.js state machine (Title→Playing→Paused→GameOver), scoring, wiring modules together
main.js bootstrap: renderer, scene, module construction, resize, the single rAF loop
```
Rules for the split:
- **Dependencies flow one way**, roughly down the list above. No circular imports. If two modules need to talk, `game.js` mediates.
- **No globals.** `CONFIG` is imported, not attached to `window`. Modules export classes or factory functions and are constructed in `main.js`.
- Each file starts with a 2–4 line header comment: what it owns and what it exports.
- Shared math helpers (clamp, lerp, damp, easing) go in `src/util.js` — not copy-pasted.
### Build in milestones, and report after each
Write the files milestone by milestone, and after each milestone print one line saying what now works. Do not write all 20 files silently and then declare victory.
- **M1 — it moves:** `game.html`, `config.js`, `util.js`, `rng.js`, `noise.js`, `world/terrain.js`, `world/streaming.js`, `player/craft.js`, `fx/camera.js`, `main.js`. A craft flies over endless streaming terrain with a chase camera. Untextured is fine here.
- **M2 — it's a game:** `input.js`, `world/props.js`, `player/collision.js`, `ui/hud.js`, `ui/screens.js`, `game.js`. Gates, orbs, shields, score, pause, game over, restart.
- **M3 — it's beautiful:** `world/sky.js`, `procedural/*`, `fx/post.js`, `fx/particles.js`, `player/trail.js`. Bloom, fog, palette, emissive neon, particles, trail.
- **M4 — it sounds right:** `audio.js`, wired into `game.js`.
The game must be runnable at the end of every milestone — never leave it in a broken intermediate state.
### The game — "NEON DRIFT"
An endless hover-racer skimming through a procedurally generated neon canyon at night.
- **World:** an infinite corridor built from streaming terrain chunks. Generate the canyon heightfield with your own fBm value/simplex noise (implement the noise yourself), carving a navigable valley down the middle whose width and curvature vary with distance. Recycle chunks ahead/behind the player — never grow the scene graph unbounded.
- **Props:** instanced crystal spires along the canyon walls, glowing boundary pylons, floating ring gates, and drifting debris. All from procedural geometry, varied by seeded noise in scale/rotation/hue.
- **Player craft:** a low-poly hovercraft assembled in code from beveled primitives, with an emissive underglow, banking animation, and a ribbon trail (a dynamically updated `BufferGeometry` strip with a fading gradient).
- **Loop:** fly forward, speed ramps up over time. Pass through ring gates for points and a combo multiplier; collect energy orbs to refill the boost meter; clipping a wall or a spire costs shield. Three shields, then game over. Score = distance × combo + gate bonuses.
- **Controls:** `A`/`D` or `←`/`→` to steer, `W`/`S` to raise/lower altitude, `Shift` or `Space` to boost, `P` to pause, `R` to restart. Also support mouse steering. Show the controls on the title screen.
- **States:** Title → Playing → Paused → Game Over → restart, all handled cleanly by one state machine in `game.js`.
### Visual direction — this must look genuinely impressive
Treat this as a showcase piece. A functional but flat-looking result is a failure.
- **Rendering setup:** `ACESFilmicToneMapping` with exposure ~1.1, `SRGBColorSpace` output, `antialias: true`.
- **Post-processing:** `EffectComposer` + `RenderPass` + `UnrealBloomPass` (strength ~0.8, radius 0.5, threshold 0.85), plus a custom `ShaderPass` doing vignette + subtle chromatic aberration + a faint film grain. Push bloom and aberration up briefly on boost and impact.
- **Sky:** a procedural gradient skydome via `ShaderMaterial` (deep indigo → magenta horizon), a `Points` starfield with per-star twinkle in the shader, and `FogExp2` whose color matches the horizon so geometry dissolves into the sky instead of popping.
- **Lighting:** cool `HemisphereLight` for ambient fill, one warm key `DirectionalLight` with soft shadows (PCFSoft), and emissive materials doing the heavy lifting for the neon look. Add a rim-light term via `onBeforeCompile` so silhouettes glow against the fog.
- **Color:** a tight, deliberate palette — cyan/magenta/violet accents on near-black terrain, with vertex-color gradients driven by height and distance. No default gray materials, no pure `0xffffff` surfaces.
- **Motion & juice:** spring-damped chase camera; FOV widening from 65° to 85° with speed; camera roll on banking; screen shake and a 60 ms hit-stop on impact; radial speed-lines and a scanline pulse during boost; additive `Points` particle bursts for gate passes and pickups; floating score popups; everything eased, nothing linear.
- **UI:** a minimal HTML/CSS overlay in a monospace font — thin glowing bars for shield and boost, score with a rolling counter, combo multiplier that scales and fades. Title and game-over cards with backdrop blur and a soft neon border. Restrained, layout-stable, no text jitter.
### Code quality
- A single `CONFIG` object in `config.js` holding every tunable (speed curve, spawn density, bloom, palette, seed) so the feel can be tuned in one place.
- Short comments explaining the non-obvious parts: noise setup, chunk recycling, the shader patches.
- Consistent naming across modules; no dead exports.
### Build priority (if you must cut scope, cut from the bottom)
1. It runs and is playable end to end, with no console errors.
2. Procedural terrain streaming + collision + scoring + game-over/restart.
3. Post-processing, sky, fog, lighting, palette.
4. Particles, trails, camera juice, HUD polish.
5. Audio.
### Definition of done — verify each before you finish
- [ ] Every relative import (`from './...'`, `from '../...'`) resolves to a file that actually exists on disk. Check this mechanically — grep every import specifier out of every source file and confirm each resolved path exists — before trusting anything else below. A single unresolved import breaks the entire module graph silently in the browser (no exception is thrown), so this check comes first.
- [ ] `python3 -m http.server` + `http://localhost:8000/game.html` plays with no console errors or warnings. Verify this by actually loading `game.html` itself through its real `<script type="module">` entry point and the real import map — in a real or headless browser — not by hand-assembling a Node.js test harness that imports individual modules with stubbed globals or mocked dependencies. A test that stubs out the module you're trying to verify proves nothing about it.
- [ ] If you verify headlessly via the Chrome DevTools Protocol, enable the `Log` and `Network` domains in addition to `Console` and `Runtime`. Broken module imports (404s, CORS failures) surface only through `Log.entryAdded` / `Network.responseReceived`, not through `Console.messageAdded` (obsolete, never fires on modern Chrome) or `Runtime.exceptionThrown` (fetch/resolution failures aren't thrown exceptions).
- [ ] No external asset file is referenced anywhere.
- [ ] No source file exceeds 250 lines; no circular imports.
- [ ] The world is endless and memory-stable; frame rate holds after minutes of play.
- [ ] Collisions, shields, scoring, pause, and restart all work.
- [ ] Bloom, fog, sky, shadows, particles, and trails are all visibly present.
- [ ] Nothing in any file is a stub.
Write the files, then finish with a short summary of the module map and anything worth tuning first.