Comparing UchenML, TensorFlow.js and LiteRT.js: Language Detection

Guesslang is the model behind @vscode/vscode-languagedetection, the library VS Code uses to auto-detect the language of untitled editors. This page runs it three ways in your browser — on UchenML compiled to WebAssembly, on upstream's TensorFlow.js engine and on LiteRT.js — and times every call as you type.

Skip the demo

Write the code below and watch three runtimes guess its language.

Type or paste code into the editor - or pick a sample. The table below lists each engine's top three guesses (expandable to five) — Uchen, upstream's TensorFlow.js and LiteRT.js — with its median time per call, measured in this tab.

Loading the language detector…

Built by Eugene Ostroukhov. How Uchen works · About & contact

Benchmarks on my machine.

Recorded on my machine: Apple M4 Max, Node v24.19.0, October 4, 2026 — 5 fresh processes per framework.

FrameworkMedian time per call [min–max]Download, uncompressed
200 chars201 KB fileCodeParameters
Uchen / WASMlangguess 0.1.0, WASM SIMD, Node9.3 µs[9.1–9.4 µs]32.9 µs[32.7–33.6 µs]408 KB959 KBuint8affine, per-tensor
Uchen / nativesame C++ Detector, arm64 native, -c opt5.5 µs[5.5–5.6 µs]15.3 µs[15.3–15.5 µs]304 KB959 KBuint8affine, per-tensor
TensorFlow.jsupstream @vscode/vscode-languagedetection 1.0.23, tfjs 3.21.0 CPU backend2,659 µs[2,615–2,737 µs]58,702 µs[57,330–60,367 µs]828 KB959 KBuint8affine, per-tensor
LiteRT.jsLiteRT.js 2.5.3, sparse .tflite, XNNPACK, 1 thread, JS featurizer30.4 µs[30.1–30.9 µs]71.2 µs[70.7–71.9 µs]9,670 KB2,703 KBfloat32not quantized
How this was measured (15 notes)
  1. Each framework ran in its own fresh process, one at a time, in an order rotated per run; each time is the median of 5 per-process medians, and the bracket is their range.
  2. Node figures are warm medians: Uchen and LiteRT.js over 1,000 batches of ≥ 2,000 µs after 50 warm-up calls; TensorFlow.js over 100 single calls after 10 warm-up calls. Real sporadic calls, TensorFlow.js above all, are slower.
  3. Native timings: raw UTF-8 bytes of the file prefix, no JS string preprocessing; the Detector reads at most the first 10.001 KB; here 0.200 / 201.420 KB of the same text as UTF-8.
  4. LiteRT.js sparse chosen as the faster variant (sum of medians 101.5 µs vs dense 204.1 µs); runtime litert_wasm_internal, XNNPACK does not take every op.
  5. LiteRT.js sparse: the bench repeats one fixed input, so no per-call input resize (when the number of distinct byte pairs changes, as in real use, the public API costs +8–9 µs at 200 chars and +5–6 µs at 10,001 chars; separate re-measurement) — this favors LiteRT.js.
  6. LiteRT.js code is the relaxed-SIMD build; the SIMD-only compat build it falls back to is 9,607.427 KB (litert_wasm_compat_internal.js 273.246 KB, litert_wasm_compat_internal.wasm 9,334.181 KB).
  7. TensorFlow.js is timed and sized as the published package: its dist/lib/index.js, loading 979.js (the CPU backend).
  8. Parameter format is how the 674,604 weights are stored. The TF.js files use uint8 affine, per-tensor quantization: each weight is one byte q, read back as min + q × scale, with min and scale stored once for each of the 9 tensors. Uchen and TensorFlow.js both expand them to float32 when loading. The .tflite that LiteRT.js runs stores them as float32, not quantized, so it is 2.8× the size of the TF.js files.
  9. Sizes are uncompressed KB (1,000 bytes) of the files as built or published; dist/lib/index.js embeds the WASM as base64; the native binary is the whole bench executable, stripped (unstripped 383.472 KB). model.json and the .bin are byte-identical to the published package's.
  10. Parity: every framework's top language matches Uchen / WASM's on every input; largest confidence difference from it 5.2e-8.
  11. Host load average (1/5/15 min): 2.52 / 9.49 / 11.15 at start, 4.19 / 5.15 / 8.42 at end; all 16 CPUs 11.3% busy over the run, the bench's own single thread included.
  12. Uchen / WASM files — code: dist/lib/index.js 407.679 KB; parameters: model.json 242.642 KB, group1-shard1of1.bin 715.908 KB.
  13. Uchen / native files — code: native_bench (arm64, stripped) 303.520 KB; parameters: model.json 242.642 KB, group1-shard1of1.bin 715.908 KB.
  14. TensorFlow.js files — code: dist/lib/index.js 705.965 KB, dist/lib/979.js 122.231 KB; parameters: model.json 242.642 KB, group1-shard1of1.bin 715.908 KB.
  15. LiteRT.js files — code: litert-race.js 28.647 KB, litert/litert_wasm_internal.js 273.035 KB, litert/litert_wasm_internal.wasm 9,367.934 KB; parameters: guesslang_sparse.tflite 2,702.260 KB, labels.json 0.393 KB.

The native build is here only because I was curious how much WASM costs over machine code. Same C++, compiled for ARM Neon instead of WASM SIMD.

TensorFlow.js is the outlier - the model is small, so most of a call runs a general-purpose graph of string operations that split the text into byte pairs, hash them and add up the results, one JavaScript kernel at a time. Uchen does the same step as one short loop of compiled C++ over the first 10,000 byte pairs, then gets on with the model's own math.

LiteRT.js is Uchen's real peer: the same general approach, optimized kernels compiled to WASM SIMD. It runs only the numeric layers; its byte pairs are hashed in JavaScript before every call. It stays close on short input, and on the big file it falls behind mostly inside those numeric layers, not in the JavaScript. The bigger difference is flexibility: Uchen can't yet define a model from JavaScript — the graph is compiled into the C++.

The sizes follow from what each one ships. Uchen's single file is built for this one model. Nothing stops it from holding more, and they would share the code already there. TensorFlow.js and LiteRT.js ship general engines that run any model, and LiteRT's engine is most of its 9,670 KB. Uchen and TensorFlow.js can read the same uint8-quantized weights, while I had to convert the weights to fp32 for LiteRT, which made it almost three times the size of the TF.js files.

Ported in one prompt.

Moving the detector onto Uchen took a single prompt:

The prompt, verbatim
/goal Build a drop-in replacement NPM module for https://github.com/microsoft/vscode-languagedetection.git based on Uchen.

Uchen: @../uchen-core/ Use Dots @../../dots/ as an example.
1. Use Bazel with submodules (like Dots).
2. Fan out agents.
3. Do as little changes to Uchen as possible - new layers and such should stay in the new module to prove Uchen is extensible.
4. Add support for loading TFlow parameters from WASM, do not preconvert.

@"systems-architect (agent)" makes the decisions. Work autonomously.

Below is the network definition the agents came up with. It is a full copy of the original one, so the parameters work as is. As usual, it is a constexpr value with the forward pass configured at compile time. I did not train this model — the weights were downloaded from the VS Code GitHub repository.

cc/model/model.h
inline constexpr size_t kEmbedding = 70;
inline constexpr size_t kHidden0 = 512;
inline constexpr size_t kHidden1 = 32;
inline constexpr size_t kParameterCount = 674604;

inline constexpr auto kModel =
    uchen::layers::FloatModel<kBuckets> | uchen::layers::Fork<2> |
    uchen::layers::Parallel(
        uchen::layers::Linear<kClasses>,
        layers::EmbeddingBagMean<kEmbedding> | uchen::layers::Linear<kHidden0> |
            uchen::layers::Relu | uchen::layers::Linear<kHidden1> |
            uchen::layers::Relu | uchen::layers::Linear<kClasses>) |
    uchen::layers::Join | uchen::layers::Reduce<uchen::PlusOp, 2> |
    uchen::layers::Logits;

static_assert(kModel.all_parameters_count() == kParameterCount);

What's running in the editor.

Model
guesslang's wide-and-deep classifier: input hashed into 5,000 byte-bigram buckets, feeding a linear branch and a 70-d embedding → 512 → 32, out to 54 classes. 675K parameters.
Weights
The same model.json and group1-shard1of1.bin, byte-identical to the npm package. One fetch, shared by Uchen and TensorFlow.js — the live table runs both on the same bytes. LiteRT.js can't run the string ops, so it gets a float32 .tflite of just the numeric layers, which I converted from those files, and hashes byte pairs in JavaScript.
Runtime
C++20 compiled to WebAssembly SIMD, shipped as a single file. The live table also loads TensorFlow.js 3.21 on its CPU backend, which is how upstream configures it, and LiteRT.js 2.5.3 on XNNPACK, single-threaded.
Parity
Same top language as upstream on all 201 non-empty cases of a 204-case golden set. Maximum confidence difference: 3e-4.
Caveats
Needs WebAssembly SIMD. Inference runs on the main thread.

Provenance

The sizes are the Size column in your network panel, not Transferred, which depends on compression. The two model files load with the page; the editor, TensorFlow.js and LiteRT.js files follow in the background, and the test file loads only when you pick its sample. To check, hash the model files against @vscode/vscode-languagedetection 1.0.23 from npm.

langguess-demo.js: 422.180 KB
model/model.json: 242.642 KB
model/model.json sha256: 100ce176367e7311e37ced0695057452991a8692029a79340a25e622893e7983
model/group1-shard1of1.bin: 715.908 KB
model/group1-shard1of1.bin sha256: fab6442698f64d5b1d2df052061d12bafd570330556819d29f48c7bcbb5889f7
in the background — langguess-editor.js: 536.955 KB
in the background — tfjs-race.js: 522.781 KB
in the background — litert-race.js: 28.647 KB
in the background — litert/litert_wasm_internal.js: 273.035 KB
in the background — litert/litert_wasm_internal.wasm: 9,367.934 KB
in the background — litert/guesslang_sparse.tflite: 2,702.260 KB
on request — large.ts.txt: 201.420 KB

Whose model this is.

The model is guesslang, by Y. Somda. The TF.js files and the API come from @vscode/vscode-languagedetection, MIT-licensed, by Microsoft.

The TensorFlow.js column runs tfjs-race.js, upstream's MIT code: their engine, not a reimplementation of it. The LiteRT.js column bundles Google's Apache-2.0 @litertjs/core.

Nothing here is retrained. Langguess is a second engine for the same weights.

One list, and it's quiet.

A new demo when there's one to try, and the day the framework is something you can actually build. Nothing else.

Double opt-in — I'll send a confirmation email first. Unsubscribe in one click.