← Back to dispatches

Teaching LLMs to Read Assembly: Idiomatic Decompilation for Modern Languages

systems-programmingreverse-engineeringinference-optimization

Based on the abstract and my knowledge of this research area, here’s the explainer:


Why Decompiling Modern Languages Is Hard—And Why It Matters

Reverse engineering has a dirty secret: most decompilers are stuck in 1995. Tools like Ghidra and IDA Pro do a reasonable job recovering C-like pseudocode from binaries, but the output is often a tangle of pointer arithmetic and goto statements that bears little resemblance to what a developer actually wrote. When the original source was a modern language like Dart or Swift, the gap grows even wider—the idioms, the garbage-collected memory model, the async primitives, and the type system all evaporate during compilation, leaving behind assembly that’s structurally alien to what the developer wrote.

This matters in security research, malware analysis, and mobile app auditing. Dart specifically powers Flutter, which means a growing fraction of production mobile and desktop apps—often handling payments, authentication, and sensitive user data—are compiled Dart binaries sitting in app stores right now, largely opaque to static analysis.

The paper takes on this problem directly: can small, fine-tuned LLMs act as idiomatic decompilers, recovering not just semantically equivalent code but code that looks like a Dart developer actually wrote it?

The Idiomatic Decompilation Goal

The distinction between “semantically correct” and “idiomatic” decompilation is significant. A traditional decompiler might recover the right computation, but express it as nested conditional branches operating on raw integer offsets. An idiomatic decompiler would ideally recover a switch on an enum, a ?. null-safe accessor, or a Future-based async chain—constructs that are meaningful to a developer trying to understand intent, not just behavior.

This framing shifts the problem. It’s not just about instruction semantics; it’s about recovering the programmer’s abstraction layer. That’s a task where LLMs, trained on enormous amounts of idiomatic source code, have a plausible structural advantage over rule-based decompilers.

Small Specialized Models Over Large Generalists

Rather than prompting GPT-4 or Claude and calling it done, the paper investigates small specialized LLMs—models fine-tuned specifically on the assembly-to-Dart translation task. This is a meaningful design choice. Large frontier models have broad coverage but limited depth on niche compilation targets; Dart on x86-64 represents a narrow enough domain that a purpose-built smaller model may outperform a general one that’s seen relatively little Dart assembly in training.

The “small” qualifier also has practical implications for deployment in security tooling, where researchers often need offline or air-gapped analysis environments where calling a cloud API isn’t an option.

Synthetic Data Augmentation

One of the paper’s core contributions is the investigation of synthetic training data. The challenge with decompilation datasets is that you need matched pairs: assembly and the corresponding high-level source. For C, decades of open-source software compiled with gcc -O2 makes this tractable. For Dart, the corpus is thinner.

The paper explores using synthetic same-language examples—likely generating Dart code programmatically or via LLM synthesis, compiling it to assembly, and using the resulting pairs to augment training data. This is then compared against training on human-written Dart examples.

This comparison is technically interesting because synthetic and human-written code differ in important ways: synthetic code may be structurally simpler, avoid unusual idioms, and have more predictable compilation patterns. Human-written code includes the full messiness of real software—complex control flow, library calls, patterns that appear in production apps. Understanding which regime produces better fine-tuned decompilers tells us something useful about the nature of the task.

The Dart-Specific Challenge

Dart compiles to native x86-64 via AOT (Ahead-Of-Time) compilation in Flutter’s production builds. The Dart runtime introduces its own calling conventions, object layouts, and GC write barriers that don’t map cleanly onto C idioms. Features like null safety (the ? type system), async/await compiled down to state machines, and Dart’s class hierarchy with vtable dispatch all leave fingerprints in assembly that a Dart-aware model needs to recognize and reverse.

This is where idiomatic decompilation becomes especially challenging: recovering the fact that a function was originally Future<String> returning with await semantics requires the model to recognize the state machine pattern the compiler emits, not just the local instruction sequence.

What to Watch For

This work opens a few threads worth following. First, the synthetic-vs-human data finding will likely inform how the field approaches other underrepresented languages (Swift, Kotlin Native, Rust)—languages increasingly present in production binaries but poorly served by existing decompilers. If synthetic augmentation works well, the bootstrapping problem for new language targets becomes much more tractable.

Second, “idiomatic” decompilation as an evaluation target implies different metrics than pure semantic equivalence. How the paper operationalizes and measures idiomaticity—BLEU scores, human evaluation, semantic similarity—will matter for how future work in this space is benchmarked.

Finally, the practical ceiling for small fine-tuned models on this task will be informative. If a 7B-parameter model fine-tuned on Dart assembly outperforms zero-shot frontier models, that’s a strong signal that specialization beats scale for structured translation tasks—a result with broad implications for security tooling pipelines.

Generated by claude-sonnet-4-6