Glyph
Glyph is a small systems language that transpiles to readable GNU C. The
compiler also embeds NeuralScript (.ns), a typed tensor DSL with CPU/CUDA AOT
code generation.
Start here
- Read the Glyphc philosophy.
- Install
glyphc. - Read the English language reference.
- Run
glyphc new helloandcd hello && glyphc build. - For tensor graphs, read the NeuralScript guide.
The project is MIT-licensed and the compiler, LSP, tests, examples, and release artifacts live in the same repository.
Философия Glyphc
Glyphc соединяет системное программирование и машинное обучение в одном пути: от читаемого исходника к нативному C/CUDA-артефакту, который можно проверить, отладить и изменить.
Glyphc — это не «ещё один язык» и не попытка заменить Python, PyTorch или LLVM одним универсальным инструментом. Это единый toolchain для двух связанных, но специализированных языков:
- Glyph (
.glyph) — системный язык с явным владением памятью, статической проверкой и читаемым C-эмитом; - NeuralScript (
.ns) — типизированный DSL для описания небольших статически известных вычислительных графов с CPU/CUDA AOT-кодогенерацией; glyphc— общий компилятор, CLI, LSP, тестовый runner и distribution point.
Сейчас общий toolchain и общая инфраструктура уже реальны. Полностью единый IR для обоих языков — направление развития, а не утверждение о текущей реализации.
Зачем существует Glyphc
Python удобен для экспериментов, но часто скрывает стоимость memory layout, dispatch, запусков и GPU-операций. C/C++ дают контроль над железом, но могут превращаться в хрупкий, слабо проверяемый слой вокруг ML-графа.
Glyphc строит мост между этими мирами:
- программист может оставаться рядом с hardware и memory layout;
- модель описывается как типизированный граф, а не как набор строковых вызовов;
- Python runtime и PyTorch bindings не являются обязательными;
- результат можно прочитать, скомпилировать обычными инструментами и встроить в C/C++-приложение.
Мы не обещаем убрать Python, CUDA, libc или системные зависимости. Мы убираем необходимость в тяжёлом промежуточном слое, который мешает видеть исходный код, сгенерированный код и поведение железа.
Четыре фундаментальных принципа
1. AI-First и читаемость человеком
AI-first означает структуру, удобную для машинного чтения, а не гарантию,
что LLM никогда не ошибётся.
В проекте для этого используются:
- явные маркеры
@module,@fn,@struct,@enum; - фиксированная структура
network,layer,forward,trainиgrad; - контракты-клозы
#guard; - машинно-читаемые diagnostics с
line:column; - детерминированный codegen;
- отсутствие скрытой семантики там, где достаточно явной конструкции.
Эти элементы создают attention anchors для человека и LLM: модель видит границы функций, модулей, типов и вычислительных этапов вместо того, чтобы угадывать намерение по свободному тексту.
Правильная формулировка:
Glyphc оптимизирован для AI-assisted development, но correctness всё равно проверяется компилятором, тестами и человеком.
2. Hardware-First и Zero-GC
В Glyph нет tracing garbage collector. Владение памятью и освобождение
ресурсов выражены явно, а flow-sensitive анализ отклоняет некоторые опасные
операции вроде повторного free() или использования освобождённого значения до
переприсваивания.
Текущая модель включает:
- явное освобождение;
drop(box);- refcounted buffers для
ListиMap; - статический анализ потока использования;
- отсутствие обязательного GC runtime.
Полные scope-based RAII и move semantics пока не реализованы в объёме, достаточном, чтобы заявлять полную совместимость с ownership-моделью Rust. Это направление развития, а не скрытая функция.
AOT-пайплайн идёт через читаемый C/C++/CUDA-код. Это позволяет приблизиться
к производительности сгенерированного C, но не является доказательством
1.0x–1.05x от «физических возможностей железа». Такие утверждения должны
подтверждаться benchmark-протоколом на конкретной платформе.
3. ML-First: тензоры и autograd как типизированные данные
В NeuralScript тензор — это не строка и не безымянный массив. У него есть ранг, размерности, dtype и проверяемые связи между операциями.
Компилятор проверяет:
- совместимость формы операндов;
- ранг
cross_entropyи classifier heads; - соответствие входов и выходов network;
- размерности embedding, attention, normalization и dense layers;
- ошибки до запуска CUDA-кода.
В проекте есть символическое связывание и разрешение dimension aliases. Это полезный shape-механизм, но пока не следует называть его полной реализацией Union-Find для произвольных constraint-графов.
Автоматическая дифференциация задаётся блоками grad { ... } (не #grad),
а fusion pass объединяет поддерживаемые группы вроде GEMM + activation +
normalization. Это ограниченный, наблюдаемый codegen-pass, а не обещание
автоматически оптимизировать любой граф.
4. Автономность deployment без ложного zero-dependency
Мы не называем систему полностью независимой от ОС: libc, pthread, компилятор, CUDA runtime и GPU driver по-прежнему важны.
Что действительно можно сказать:
- Python runtime не обязателен;
- PyTorch bindings не обязательны;
- tracing GC не обязателен;
- сгенерированный C/CUDA-код можно статически слинковать;
- C-ABI (
ns_runtime.h) позволяет встроить результат в существующее C/C++-приложение.
Поддержка автономного inference 27B-моделей, cold-start 1.5–3 секунды и динамической библиотеки в один файл пока являются целями развития, а не текущими возможностями v2.1.0.
Что Glyphc сознательно не обещает
| Формулировка | Корректная позиция проекта |
|---|---|
| Zero-GC | Да для Glyph: нет tracing GC |
| RAII и move semantics | Частичная модель и дорожная карта |
| Производительность C/Rust | Цель — измеримо приблизиться к C-эмиту; гарантий нет |
| Symbolic Union-Find | Есть symbolic dimension binding; полный constraint solver не заявлен |
#grad | Реальный синтаксис — grad { ... } |
| Inference 27B | Не заявлен и не поддерживается как стабильная функция |
| 1.5–3 секунды cold start | Не подтверждённый benchmark |
| One binary | Да для glyphc; CUDA всё равно требует CUDA runtime/driver |
| Zero dependencies | Нет обязательных Python/GC-зависимостей, но есть системные toolchain-зависимости |
| AI без галлюцинаций | AI-readable structure, а не гарантия корректности LLM |
| Full kernel fusion | Есть ограниченный fusion pass, не весь граф |
Эта таблица — часть дизайна, а не юридический дисклеймер. Она защищает проект от обещаний, которые невозможно проверить.
Сильные стороны
- Уникальное сочетание: системный язык и специализированный AOT ML DSL используют один toolchain.
- Инспектируемость: C/CUDA-результат остаётся исходным кодом, который можно прочитать и изменить.
- Явная память: нет обязательного GC и неявных пауз, за которые отвечает runtime.
- Статическая проверка форм: ошибки neural graph обнаруживаются до запуска тяжёлого обучения.
- C/C++ интеграция: generated runtime можно использовать как C-ABI внутри существующей системы.
- AI-friendly структура: строгие маркеры и стабильные diagnostics помогают LLM и инструментам анализа.
- Воспроизводимость: sorted generic/constant emission и regression tests устраняют случайный порядок функций.
Слабые стороны
- Scope шире, чем у большинства языков: systems и ML требуют разных специалистов и длинной дорожной карты.
- NeuralScript пока не является полной заменой PyTorch: у него ограниченный набор операторов и статических форм.
- Native fp16, ROCm и Metal ещё не являются полноценными production backend’ами.
- C backend ориентирован на POSIX/GNU C; Windows binary release есть, но native Win32 runtime требует отдельной работы.
- Экосистема, package manager и Marketplace-расширение ещё находятся на ранней стадии.
- Производительность не должна заменяться маркетинговыми числами: нужны воспроизводимые hardware-specific benchmarks.
Формула продукта
Пиши системно. Описывай вычисления. Компилируй прозрачно.
English:
Write systems. Describe tensors. Compile both to native artifacts you can inspect.
Glyphc не пытается быть Python, PyTorch, LLVM и Rust одновременно. Его уникальная роль — нативный AI без обязательного Python-слоя, с типизированными графами и инспектируемым C/CUDA-результатом.
Направление развития
- Стабилизировать memory model и опубликовать crate в crates.io.
- Сделать NNS numerical test suite и расширить поддерживаемые операции.
- Реализовать настоящий fp16/bf16 path либо окончательно обозначить его как experimental.
- Добавить NNS LSP diagnostics, package management и VS Code Marketplace.
- Расширить shape inference, включая rank-3+ случаи.
- Принять решение по native ROCm и Metal, не выдавая placeholders за полноценные backend’ы.
- Публиковать benchmark artifacts вместе с CPU, compiler и CUDA versions.
Glyph Language Reference
Status: v2.1.0. This document describes the actually working syntax
(verified against examples/ and tests). Undocumented constructs are not considered part of the language.
Table of Contents
- File Structure
- Types
- Variables and Constants
- Functions
- Operators
- Control Flow
- Strings and Lists
- Structs and Methods
- Enums
- Modules
- Tests and Asserts
- Concurrency
- Standard Library
1. File Structure
A file is a sequence of top-level declarations:
@module examples.demo // top-level module name
@use other.module; // import another module
@const SCALE: Float64 = 1.5;
@struct Point { x: Float64, y: Float64 }
@enum Op { Plus, Minus, MulDiv(Float64, Float64) }
@fn twice(x: Int64) -> Int64 { return x * 2; }
@test @fn test_twice() -> Void { assert_eq(twice(2), 4, "2 * 2 == 4"); }
Attributes (written with @):
| Attribute | Applies to | Purpose |
|---|---|---|
@module | file | module name |
@use | file | module import |
@fn | function | function declaration |
@struct | struct | struct declaration |
@enum | enum | enum declaration |
@const | constant | named constant |
@pub | fn/struct/enum | public visibility (for imported modules) |
@test | fn | test function (see docs/testing.md) |
Comments: // line, /* block */.
2. Types
| Type | C representation | Notes |
|---|---|---|
Int64 | int64_t | signed integer |
UInt64 | uint64_t | unsigned |
Float64 | double | floating point |
Bool | int | true/false |
String | char* | strings |
Bytes | char* | raw bytes |
List<T> | GlyphList (refcounted fat-struct) | literal [..], range .., .len(), .append(), .free(), slices, for x in xs, ++, == for POD |
Option<T> | GlyphBox* (C: void*) | box {tag, data}; Some(x)/None |
Result<T, E> | GlyphBox* (C: void*) | box {tag, data}; Ok(x)/Err(e) |
&T | T* | reference to a value (for enums in match) |
MyStruct, MyEnum | struct/tagged union | user-defined types |
Result / Option with payloads
Result<Ok, E> and Option<T> are built from variants and matched in match.
Payload is copied to the heap (box { tag, data }); tag 1 = Ok/Some,
0 = Err/None:
@fn divide(a: Int64, b: Int64) -> Result<Int64, String> {
if b == 0 {
return Result::Err("division by zero");
}
return Result::Ok(a / b);
}
@fn main() {
let q: Int64 = match divide(10, 2) {
| Ok(v) => v,
| Err(e) => -1
};
print_int(q);
let parsed: Option<Float64> = parse_float("2.5");
let f: Float64 = match parsed {
| Some(v) => v,
| None => 0.0
};
print_float(f);
}
Notes:
Option::None/Result::Ok(x)without context yields an unknown parameter (Type::Void); inletit unifies with the declared type (let n: Option<Int64> = Option::None)parse_int/parse_float/to_int/to_floatreturn boxedOption(Some/None) suitable formatch
Generic functions
Type parameters are declared in angle brackets after the function name and can
be used in the signature and body, including nested (List<T>, Option<T>,
Result<T, E>). The compiler monomorphizes each instantiation into a
separate C function (identity_i64, first_str_i64):
@fn identity<T>(x: T) -> T {
return x;
}
@fn first<T, K>(a: T, b: K) -> T {
return a;
}
@fn main() {
let a: Int64 = identity(5); // T = Int64 from argument
let s: String = first("hi", 42); // T = String, K = Int64
print_int(a);
}
Type inference:
- from call argument types (including nested:
List<Int64>givesT = Int64); - from
letannotation for parameters appearing only in the return type (let r: Result<String, Int64> = ok_wrap("fine")infersE = Int64).
Limitations in v2.1: only functions are generic (structs/enums/impl are
concrete); explicit type arguments (f::<Int64>) are not supported;
calling a generic function inside a generic body requires concrete types
(otherwise — error Cannot infer type arguments); type [T; N] is described
but the literal always produces List<T>, so arrays are not constructible.
Literals:
42 // Int64
498_351_8 // digit separators
3.14 // Float64
true false // Bool
"hello" // String
0xFF // hex
3. Variables and Constants
let x: Int64 = 42; // immutable
let mut acc: Float64 = 0.0; // mutable
let y = 42; // untyped let: type inferred (Int64, String, List<T>, ...)
let r = 0..5; // List<Int64]
let s = "a" ++ "b"; // String
x = 10; // ❌ type error: x is immutable
acc = acc + 1; // ✅
@const PI: Float64 = 3.141592653589793;
@const MAX_RETRIES: Int64 = 3;
Untyped let infers the type from the value (literals, List, Range,
concatenation, calls). Unknown type is a codegen error
(cannot infer type of untyped let), not garbage in C.
Constants compile to static const expressions and may reference other
constants and literals.
4. Functions
@fn add(a: Int64, b: Int64) -> Int64 { return a + b; }
@fn log_message(msg: String) -> Void { println(msg); }
@fn default() -> Int64 { 0 } // expression body
@fn make() -> Float64 { return 2.0; }
- Parameters are typed; return type is indicated via
->,Voidmeans no value. - Arguments are passed by value; references
&Tare available. - Unqualified calls inside the same module work as-is.
@fn async(and@fn asyncmethods in@impl) returns a lazy handle — see “Concurrency” below. Async-generic functions are not supported.
5. Operators
Arithmetic: + - * / %.
Comparisons: == != < > <= >=. Logic: && || !.
String concatenation: ++ — "Hello, " ++ "World!". When concatenating
with numbers, use conversion: "x=" ++ int_to_string(x).
Type cast as: let f: Float64 = 42 as Float64;.
Reference &: let r: &Shape = &circle; (used for match by value
without copying).
6. Control Flow
if / else
if x > 0 {
print("positive");
} else if x < 0 {
print("negative");
} else {
print("zero");
}
if is also an expression: when both branches evaluate to a value of the same
scalar type (Int64, UInt64, Float64, Bool), the whole if/else has that
type and can be assigned. An integer literal branch promotes to a sibling’s
Float64/UInt64 type. Heap-typed branches (String, List<T>, structs) are
not supported as if expression values.
let hi: Int64 = if coins > 100 { 2; } else { 1; };
let scale: Float64 = if fast { 1.5; } else { 1; };
while
let i: Int64 = 0;
while i < 5 {
i = i + 1;
}
loop / break / continue
let i: Int64 = 0;
loop {
i = i + 1;
if i >= 100 { break; }
}
for .. in (integer range)
let sum: Int64 = 0;
for i in 0..5 { // 0..5 — excluding 5 → sum 10
sum = sum + i;
}
for j in 1..=5 { // 1..=5 — inclusive
sum = sum + j;
}
match
match value {
| 1 => print("one")
| 2 => print("two")
| _ => print("other")
}
match status {
| HttpStatus::Ok => "OK"
| HttpStatus::NotFound => "Not Found"
| _ => "Unknown"
}
match shape {
| Shape::Circle(r) => 3.14159 * r * r
| Shape::Rectangle(w, h) => w * h
| Shape::Point => 0.0
}
Patterns: literals, Enum::Variant, Enum::Variant(bindings), _ (wildcard).
Arm bodies are expressions (including calls terminated with ;).
guard
@fn divide(a: Float64, b: Float64) -> DivisionResult {
#guard(b != 0.0) else {
return DivisionResult::DivisionByZero;
};
return DivisionResult::Success(a / b);
}
#guard(condition) else { ... }; — if the condition is false, the else
block executes and the current function returns; ; after else is mandatory.
7. Strings and Lists
let greeting: String = "Hello, " ++ "Glyph!";
let n: Int64 = len(greeting); // 13
let sub: String = substring(greeting, 0, 5); // "Hello"
let has: Bool = contains(greeting, "Glyph"); // true
let upper: String = to_upper(greeting); // "HELLO, GLYPH!"
let words: String = split("a,b,c", ","); // CSV fragment
let num: Option<Int64> = parse_int("42"); // Option<Int64>
let s: String = int_to_string(42);
let arr: List<Int64> = [10, 20, 30]; // list literal
print_int(arr[1]); // 20 — indexing
let points: List<P> = [P { x: 1.0, y: 2.0 }];
print_float(points[0].y); // 2.0 — struct indexing
let range: List<Int64> = 0..10; // [0, 1, ..., 9]
print_int(range.len()); // 10 — length
range.append(10); // buffer growth
let both: List<Int64> = arr ++ range; // list concatenation
for x in arr { // iteration
print_int(x);
}
Lists are fat-structs {data, len, elem_size, refs} over a heap buffer:
literals and ranges allocate, append grows the buffer, indexing arr[i]
returns element T. Copying a list (let b = a, passing to a function)
bumps the refcount: the buffer lives while at least one copy exists, so
a.free(); b[0] is safe. append with multiple owners does copy-on-write
(separate buffer for this owner; append inside a function is not visible
outside). The element type of a literal is inferred from the first element.
let without annotation also infers the type (let xs = [1, 2] gives List<Int64>).
Slices xs[a..b] / xs[a..=b] return a new List<T> (range copy, bounds
clamped). POD-scalar lists (Int64, UInt64, Float64, Bool) compare by
value (==/!= via len + memcmp); others are a loud codegen error.
xs.free() releases one ref and zeroes the list (buffer stays while other
copies exist).
7.1. Map
Map<String, V> — chained hash map (FNV-1a, key always String), values
copied by elem_size(V). Literal #{ "key": value, ... } builds a map;
value type is inferred from the first pair (empty → Map<String, Void>).
@fn main() -> Void {
let m: Map<String, Int64> = #{ "alice": 90, "bob": 75 };
print_int(m["alice"]); // read; missing key — abort
m["carol"] = 60; // write via index
m.put("bob", 80); // write via method
let got: Option<Int64> = m.get("dave"); // Option instead of abort
let v: Int64 = match got {
| Some(x) => x
| None => 0
};
drop(got); // Option box requires drop(box)
print_int(m.len());
for k in m { // iteration yields keys (String)
print(k);
}
m.free(); // free storage
}
Note: m["k"] on a missing key aborts — for checked access use m.get(k)
(returns Option<V>; Some box is freed via drop(box)). Iteration
for k in m walks bucket chains (order = hash, not insertion order) and yields
keys; values are read via m[k].
7.2. Static use-after-free protection
The compiler tracks v.free() calls for List and Map in statement flow
and rejects subsequent use of v until reassignment:
let xs: List<Int64> = [1, 2, 3];
xs.free();
xs.append(4); // ERROR: use after free: 'xs'
xs.free(); // ERROR: double free: 'xs'
Legal path — reassign after free() (slot revives):
xs.free();
xs = [4, 5]; // reassignment resets the flag
xs.append(6);
Analyzer limitations: statement-flow, per-function — does not track aliases
(let y = xs; xs.free(); y.len() is not caught), and free() inside a
match arm is treated as executed after the match. Coverage: reads,
indexing, iteration, argument passing, and repeated .free().
8. Structs and Methods
@struct Point {
x: Float64,
y: Float64,
}
let p: Point = Point { x: 0.0, y: 0.0 };
let dx: Float64 = p2.x - p1.x;
Methods via @impl
@impl <Type> { @fn ... } declares methods for a type. The receiver is
passed as the first parameter; call obj.method(a, b) desugars to
Type_method(obj, a, b).
@impl Point {
@fn norm(p: Point) -> Float64 {
return sqrt(p.x * p.x + p.y * p.y);
}
@fn scaled(p: Point, k: Float64) -> Point {
return Point { x: p.x * k, y: p.y * k };
}
}
let p: Point = Point { x: 3.0, y: 4.0 };
let n: Float64 = p.norm(); // → Point_norm(p)
let q: Point = p.scaled(2.0); // → Point_scaled(p, 2.0)
Rules:
- first method parameter must have type (or
&-reference to)Type— same as the call object; - remaining parameters are regular arguments;
args.len() == params.len() - 1; - method body accesses receiver fields via the parameter (
p.x), noself; - methods from other modules are not prefixed (§10), method names are
Type_method; methods cannot be imported as functions.
9. Enums
@enum Color { Red, Green, Blue } // unit variants
@enum Shape {
Circle(Float64), // variant with data
Rectangle(Float64, Float64),
Point,
}
@enum HttpStatus { // discriminants
Ok = 200,
NotFound = 404,
}
Construction: Shape::Circle(5.0), HttpStatus::Ok, Color::Red.
Matching:
match s {
| Shape::Circle(r) => 3.14159 * r * r
| Shape::Rectangle(w, h) => w * h
| Shape::Point => 0.0
}
:: parsing rule: a path segment starting with uppercase is an enum
constructor (Shape::Circle(...)); starting with lowercase — a qualified
module function call (math::add(...)).
10. Modules
// math.glyph —— module
@module my.math
@pub @fn add(a: Int64, b: Int64) -> Int64 { return a + b; }
// main.glyph —— entry
@use my.math;
@fn main() -> Void {
let sum: Int64 = my.math::add(1, 2);
}
- Entry file (the one passed to
glyphc run/compile) is not prefixed. - Imported module names are prefixed (
add→my_math_add); references to own functions inside the module are rewritten automatically. @pub— public visibility; by default symbols are private, but in v2.1 the marking is not a hard restriction.
11. Tests and Asserts
See docs/testing.md.
12. Concurrency
Calling @fn async does not execute the body but returns a lazy handle Async<T>.
spawn h; enqueues the task in the shared FIFO worker pool (idempotent);
the pool is M:N (M tasks on N threads), N = core count, capped at 8,
overridable via GLYPH_WORKERS=K. h await waits for the result: executes
synchronously if the handle has not yet been launched; otherwise waits on a
condvar. The waiting thread “helps” the pool — while the target is not ready
it drains the queue and executes other tasks (work-sharing), so nested await
inside async functions does not deadlock the pool. Repeated await returns the
cached value:
@fn async fetch(url: String) -> String {
return url;
}
@fn main() {
let h: Async<String> = fetch("http://x");
spawn h;
println(h await);
}
List/Map across worker boundaries are passed by value-struct:
the trampoline puts the struct in the cell, and each await deep-copies it
(GlyphList/GlyphMap memcpy data + fresh refcount). So the handle keeps a
private result instance: mutation and free() of the value obtained via
await are safe and only corrupt that copy; repeated await yields the
original value again. (Handle result cell is not freed — same leak class as
Int64/String: runtime without GC.)
Typed channels are created with Channel<T>(capacity):
capacity 0 — rendezvous (direct handoff without buffer), > 0 — buffered.
send(v) blocks when the buffer is full and returns false if the channel is closed;
recv() blocks when empty and returns None when the channel is closed
and empty; close() wakes all waiters:
@fn async produce(ch: Channel<Int64>) -> Int64 {
ch.send(1);
ch.send(2);
ch.close();
return 2;
}
@fn main() {
let ch: Channel<Int64> = Channel<Int64>(4);
spawn produce(ch);
let m: Option<Int64> = ch.recv();
let v: Int64 = match m {
| Some(x) => x,
| None => -1
};
print_int(v);
}
select: waiting for the first ready event
select { ... } waits for one of several events and executes the matching
block. An arm is either | name: Type <- channel.recv(), or
| name: Type <- handle await, or | timeout(ms), or | default.
The first source in order that already has data fires;
for recv from a buffered channel and for await, task submission to the pool
happens automatically. timeout(ms) fires if no event becomes ready within
ms milliseconds; default — if no ready event exists right now.
timeout and default cannot be specified together (in each select),
and there cannot be multiple timeout arms. Waiting logic polls at ~1 ms
steps, so timeouts are honest and the worker pool keeps running:
@fn async worker(ch: Channel<Int64>) -> Int64 {
ch.send(42);
ch.close();
return 42;
}
@fn main() {
let ch: Channel<Int64> = Channel<Int64>(1);
let h: Async<Int64> = worker(ch);
select {
| v: Int64 <- ch.recv() => print_int(v),
| v: Int64 <- h await => print_int(v),
| timeout(100) => print_int(-1)
}
}
recv from a closed and drained channel is immediately ready and returns
the zero value of the type (analogous to None).
Limitations: runtime without GC (handles, channels and lists are not freed);
Option/Result boxes: temporaries (call result directly in match)
are freed automatically, named ones via drop(box); select arm must be
exactly recv() or await (arbitrary expressions including recv_timeout
are not allowed); threads are not preempted at task level — a long compute
at the task root occupies the worker entirely (via nested await the worker
switches to other tasks); async-generic functions are forbidden;
send(&ref) is forbidden by type (stack address must not be passed
between threads). List/Map values in a channel are passed as struct copies
with shared buffer (COW, like String): after send the sender must not
reuse or free the value — the receiver owns it.
13. Standard Library
The library is available without @use; signatures are declared in the
compiler, bodies in stdlib/.
print(msg: String) -> Void
println(msg: String) -> Void
eprintln(msg: String) -> Void
print_int(value: Int64) -> Void
print_float(value: Float64) -> Void
print_bool(value: Bool) -> Void
read_line() -> String
sqrt(x: Float64) -> Float64
pow(base: Float64, exp: Float64) -> Float64
abs(x: Float64) -> Float64
floor(x: Float64) -> Int64
ceil(x: Float64) -> Int64
round(x: Float64) -> Int64
min(a: Float64, b: Float64) -> Float64
max(a: Float64, b: Float64) -> Float64
clamp(value: Float64, low: Float64, high: Float64) -> Float64
log(x: Float64) -> Float64 log2(x: Float64) -> Float64
log10(x: Float64) -> Float64 sin(x: Float64) -> Float64
cos(x: Float64) -> Float64 tan(x: Float64) -> Float64
len(s: String) -> Int64
substring(s: String, start: Int64, end: Int64) -> String
contains(s: String, sub: String) -> Bool
starts_with(s: String, prefix: String) -> Bool
ends_with(s: String, suffix: String) -> Bool
replace(s: String, from: String, to: String) -> String
split(s: String, delimiter: String) -> String
trim(s: String) -> String
to_upper(s: String) -> String to_lower(s: String) -> String
char_at(s: String, index: Int64) -> String
parse_int(s: String) -> Option<Int64>
parse_float(s: String) -> Option<Float64>
int_to_string(value: Int64) -> String
float_to_string(value: Float64) -> String
bool_to_string(value: Bool) -> String
to_int(s: String) -> Option<Int64>
to_float(s: String) -> Option<Float64>
to_bool(s: String) -> Option<Bool>
file_read(path: String) -> String
file_write(path: String, content: String) -> Bool
file_append(path: String, content: String) -> Bool
file_exists(path: String) -> Bool
create_dir(path: String) -> Bool
remove_file(path: String) -> Bool
list_dir(path: String) -> String
file_copy(src: String, dst: String) -> Bool
file_rename(old: String, new: String) -> Bool
alloc(size: Int64) -> Bytes
free(ptr: Bytes) -> Void
memcpy(dst: Bytes, src: Bytes, size: Int64) -> Void
memset(dst: Bytes, value: Int64, size: Int64) -> Void
assert(condition: Bool, msg: String) -> Void
assert_eq(a: Int64|UInt64|Float64|Bool|String, b: ..., msg: String) -> Void
assert_ne(a: ..., b: ..., msg: String) -> Void
assert_true(condition: Bool, msg: String) -> Void
assert_false(condition: Bool, msg: String) -> Void
Язык Glyph — справочник
Статус: v2.1.0. Документ описывает реально работающий синтаксис
(проверено на examples/ и тестах). Незадокументированные конструкции не считаются частью языка.
1. Структура файла
Файл — последовательность объявлений верхнего уровня:
@module examples.demo // имя модуля верхнего уровня
@use other.module; // импорт другого модуля
@const SCALE: Float64 = 1.5;
@struct Point { x: Float64, y: Float64 }
@enum Op { Plus, Minus, MulDiv(Float64, Float64) }
@fn twice(x: Int64) -> Int64 { return x * 2; }
@test @fn test_twice() -> Void { assert_eq(twice(2), 4, "2 * 2 == 4"); }
Атрибуты (пишутся через @):
| Атрибут | Применяется к | Назначение |
|---|---|---|
@module | файл | имя модуля |
@use | файл | импорт модуля |
@fn | функция | объявление функции |
@struct | структура | объявление структуры |
@enum | перечисление | объявление перечисления |
@const | константа | именованная константа |
@pub | fn/struct/enum | публичная видимость (для импортируемых модулей) |
@test | fn | тестовая функция (см. docs/testing.md) |
Комментарии: // строка, /* блок */.
2. Типы
| Тип | C-представление | Примечание |
|---|---|---|
Int64 | int64_t | целые |
UInt64 | uint64_t | беззнаковые |
Float64 | double | числа с плавающей точкой |
Bool | int | true/false |
String | char* | строки |
Bytes | char* | сырые байты |
List<T> | GlyphList (refcounted fat-struct) | литерал [..], диапазон .., .len(), .append(), .free(), срезы, for x in xs, ++, == для POD |
Option<T> | GlyphBox* (в C — void*) | коробка {tag, data}; Some(x)/None |
Result<T, E> | GlyphBox* (в C — void*) | коробка {tag, data}; Ok(x)/Err(e) |
&T | T* | ссылка на значение (для перечислений в match) |
MyStruct, MyEnum | struct/tagged union | пользовательские типы |
Result / Option с данными
Result<Ok, E> и Option<T> строятся вариантами и сопоставляются в match.
Payload копируется в кучу (коробка { tag, data }); тег 1 — Ok/Some,
0 — Err/None:
@fn divide(a: Int64, b: Int64) -> Result<Int64, String> {
if b == 0 {
return Result::Err("division by zero");
}
return Result::Ok(a / b);
}
@fn main() {
let q: Int64 = match divide(10, 2) {
| Ok(v) => v,
| Err(e) => -1
};
print_int(q);
let parsed: Option<Float64> = parse_float("2.5");
let f: Float64 = match parsed {
| Some(v) => v,
| None => 0.0
};
print_float(f);
}
Замечания:
Option::None/Result::Ok(x)без контекста дают неизвестный параметр (Type::Void); вletон унифицируется с объявленным типом (let n: Option<Int64> = Option::None)parse_int/parse_float/to_int/to_floatвозвращают коробкуOption(Some/None) и пригодны дляmatch
Обобщённые функции
Параметры типа объявляются в угловых скобках после имени функции и могут
использоваться в сигнатуре и теле, в том числе вложенно (List<T>,
Option<T>, Result<T, E>). Компилятор мономорфизирует каждую инстанциацию
в отдельную C-функцию (identity_i64, first_str_i64):
@fn identity<T>(x: T) -> T {
return x;
}
@fn first<T, K>(a: T, b: K) -> T {
return a;
}
@fn main() {
let a: Int64 = identity(5); // T = Int64 из аргумента
let s: String = first("hi", 42); // T = String, K = Int64
print_int(a);
}
Вывод типов:
- по типам аргументов вызова (включая вложенные:
List<Int64>даётT = Int64); - из аннотации
letдля параметров, встречающихся только в возвращаемом типе (let r: Result<String, Int64> = ok_wrap("fine")выводитE = Int64).
Ограничения v2.1: обобщаются только функции (структуры/enum/impl —
конкретные); явные аргументы типов (f::<Int64>) не поддерживаются;
вызов generic-функции внутри generic-тела требует конкретных типов
(иначе — ошибка Cannot infer type arguments); тип [T; N] описан, но литерал всегда порождает List<T>, поэтому
массивы неконструируемы.
Литералы:
42 // Int64
498_351_8 // разделители разрядов
3.14 // Float64
true false // Bool
"hello" // String
0xFF // hex
3. Переменные и константы
let x: Int64 = 42; // неизменяемая
let mut acc: Float64 = 0.0; // изменяемая
let y = 42; // untyped let: тип выводится (Int64, String, List<T>, ...)
let r = 0..5; // List<Int64]
let s = "a" ++ "b"; // String
x = 10; // ❌ ошибка типов: x неизменяемая
acc = acc + 1; // ✅
@const PI: Float64 = 3.141592653589793;
@const MAX_RETRIES: Int64 = 3;
Untyped let выводит тип из значения (литералы, List, Range,
конкатенация, вызовы). Неизвестный тип — ошибка кодгена
(cannot infer type of untyped let), а не мусор в C.
Константы компилируются в static const-выражения и могут ссылаться
на другие константы и литералы.
4. Функции
@fn add(a: Int64, b: Int64) -> Int64 { return a + b; }
@fn log_message(msg: String) -> Void { println(msg); }
@fn default() -> Int64 { 0 } // тело-выражение
@fn make() -> Float64 { return 2.0; }
- Параметры типизированы; тип возврата указывается через
->,Void— без значения. - Аргументы передаются по значению; доступны ссылки
&T. - Неквалифицированные вызовы внутри собственного модуля работают как есть.
@fn async(и@fn asyncметоды в@impl) возвращают ленивый хендл — см. раздел «Конкурентность» ниже. Async-generic функции не поддерживаются.
5. Операторы
Арифметика: + - * / %.
Сравнения: == != < > <= >=. Логика: && || !.
Конкатенация строк: ++ — "Hello, " ++ "World!". При конкатенации
с числами используйте конвертацию: "x=" ++ int_to_string(x).
Приведение типов as: let f: Float64 = 42 as Float64;.
Ссылка &: let r: &Shape = &circle; (используется для match по значению
без копирования).
6. Управляющий поток
if / else
if x > 0 {
print("positive");
} else if x < 0 {
print("negative");
} else {
print("zero");
}
if is also an expression: when both branches evaluate to a value of the same
scalar type (Int64, UInt64, Float64, Bool), the whole if/else has that
type and can be assigned. An integer literal branch promotes to a sibling’s
Float64/UInt64 type. Heap-typed branches (String, List<T>, structs) are
not supported as if expression values.
let hi: Int64 = if coins > 100 { 2; } else { 1; };
let scale: Float64 = if fast { 1.5; } else { 1; };
while
let i: Int64 = 0;
while i < 5 {
i = i + 1;
}
loop / break / continue
let i: Int64 = 0;
loop {
i = i + 1;
if i >= 100 { break; }
}
for .. in (целочисленный диапазон)
let sum: Int64 = 0;
for i in 0..5 { // 0..5 — не включая 5 → сумма 10
sum = sum + i;
}
for j in 1..=5 { // 1..=5 — включительно
sum = sum + j;
}
match
match value {
| 1 => print("one")
| 2 => print("two")
| _ => print("other")
}
match status {
| HttpStatus::Ok => "OK"
| HttpStatus::NotFound => "Not Found"
| _ => "Unknown"
}
match shape {
| Shape::Circle(r) => 3.14159 * r * r
| Shape::Rectangle(w, h) => w * h
| Shape::Point => 0.0
}
Паттерны: литералы, Enum::Variant, Enum::Variant(bindings), _ (wildcard).
Body arms — выражения (в т.ч. вызовы, завершившиеся ;).
guard
@fn divide(a: Float64, b: Float64) -> DivisionResult {
#guard(b != 0.0) else {
return DivisionResult::DivisionByZero;
};
return DivisionResult::Success(a / b);
}
#guard(условие) else { ... }; — если условие ложно, выполняется блок else
и текущая функция возвращается; после else обязательна ;.
7. Строки и списки
let greeting: String = "Hello, " ++ "Glyph!";
let n: Int64 = len(greeting); // 13
let sub: String = substring(greeting, 0, 5); // "Hello"
let has: Bool = contains(greeting, "Glyph"); // true
let upper: String = to_upper(greeting); // "HELLO, GLYPH!"
let words: String = split("a,b,c", ","); // CSV-фрагмент
let num: Option<Int64> = parse_int("42"); // Option<Int64>
let s: String = int_to_string(42);
let arr: List<Int64> = [10, 20, 30]; // литерал списка
print_int(arr[1]); // 20 — индексация
let points: List<P> = [P { x: 1.0, y: 2.0 }];
print_float(points[0].y); // 2.0 — индексация структур
let range: List<Int64> = 0..10; // [0, 1, ..., 9]
print_int(range.len()); // 10 — длина
range.append(10); // рост буфера
let both: List<Int64> = arr ++ range; // конкатенация списков
for x in arr { // итерация по списку
print_int(x);
}
Списки — fat-struct {data, len, elem_size, refs} поверх heap-буфера:
литералы и диапазоны выделяют память, append растягивает буфер, индексация
arr[i] возвращает элемент T. Копирование списка (let b = a, передача в
функцию) считается ref-count’ом: буфер жив, пока жива хотя бы одна копия,
поэтому a.free(); b[0] безопасно. append при нескольких владельцах делает
copy-on-write (отдельный буфер для этого владельца; append внутри функции
не виден снаружи). Тип элемента литерала выводится по первому элементу.
let без аннотации тоже выводит тип (let xs = [1, 2] даёт List<Int64>).
Срезы xs[a..b] / xs[a..=b] возвращают новый List<T> (копия диапазона,
границы клампятся). Списки POD-скаляров (Int64, UInt64, Float64,
Bool) сравниваются по значению (==/!= через len + memcmp);
остальные — громкая ошибка кодгена. xs.free() отпускает один ref и
обнуляет список (буфер остаётся, пока есть другие копии).
7.1. Map
Map<String, V> — chained-хэшмап (FNV-1a, ключ всегда String), значения
копируются по elem_size(V). Литерал #{ "key": value, ... } строит map;
тип значения выводится по первой паре (пустой — Map<String, Void>).
@fn main() -> Void {
let m: Map<String, Int64> = #{ "alice": 90, "bob": 75 };
print_int(m["alice"]); // чтение; отсутствующий ключ — abort
m["carol"] = 60; // запись через индекс
m.put("bob", 80); // запись методом
let got: Option<Int64> = m.get("dave"); // Option вместо abort
let v: Int64 = match got {
| Some(x) => x
| None => 0
};
drop(got); // Option-бокс требует drop(box)
print_int(m.len());
for k in m { // итерация даёт ключи (String)
print(k);
}
m.free(); // освободить хранилище
}
Внимание: m["k"] по отсутствующему ключу аварийно завершает программу —
для проверяемого доступа используйте m.get(k) (возвращает Option<V>;
Some-бокс освобождается через drop(box)). Итерация for k in m
проходит по bucket-цепям (порядок = хэшу, не порядку вставки) и выдаёт
ключи; значения читаются через m[k].
7.2. Статическая защита от use-after-free
Компилятор отслеживает в потоке операторов вызовы v.free() для List и
Map и отклоняет последующее обращение к v до переприсваивания:
let xs: List<Int64> = [1, 2, 3];
xs.free();
xs.append(4); // ОШИБКА: use after free: 'xs'
xs.free(); // ОШИБКА: double free: 'xs'
Легальный путь — переприсвоить переменную после free() (слот «оживает»):
xs.free();
xs = [4, 5]; // переприсваивание сбрасывает флаг
xs.append(6);
Ограничения анализатора: он statement-flow и пофункциональный — не отслеживает
алиасы (let y = xs; xs.free(); y.len() не ловится), а free() внутри ветки
match трактуется как выполненный и после match. Покрываются чтение,
индексация, итерация, передача в аргументы и повторный .free().
8. Структуры и методы
@struct Point {
x: Float64,
y: Float64,
}
let p: Point = Point { x: 0.0, y: 0.0 };
let dx: Float64 = p2.x - p1.x;
Методы через @impl
@impl <Type> { @fn ... } объявляет методы типа. Ресивер передаётся
первым параметром, вызов obj.method(a, b) превращается в
Type_method(obj, a, b).
@impl Point {
@fn norm(p: Point) -> Float64 {
return sqrt(p.x * p.x + p.y * p.y);
}
@fn scaled(p: Point, k: Float64) -> Point {
return Point { x: p.x * k, y: p.y * k };
}
}
let p: Point = Point { x: 3.0, y: 4.0 };
let n: Float64 = p.norm(); // → Point_norm(p)
let q: Point = p.scaled(2.0); // → Point_scaled(p, 2.0)
Правила:
- первый параметр метода должен иметь тип (или
&-ссылку на тип)Type— тот же, что и у объекта вызова; - остальные параметры — обычные аргументы;
args.len() == params.len() - 1; - тело метода обращается к полям ресивера через параметр (
p.x), безself; - методы других модулей не префиксуются (§10), имена методов —
Type_method; методы нельзя импортировать как функции.
9. Перечисления
@enum Color { Red, Green, Blue } // unit-варианты
@enum Shape {
Circle(Float64), // вариант с данными
Rectangle(Float64, Float64),
Point,
}
@enum HttpStatus { // дискриминанты
Ok = 200,
NotFound = 404,
}
Конструкция: Shape::Circle(5.0), HttpStatus::Ok, Color::Red.
Сопоставление:
match s {
| Shape::Circle(r) => 3.14159 * r * r
| Shape::Rectangle(w, h) => w * h
| Shape::Point => 0.0
}
Правило разбора ::: отрезот патh с заглавной буквы — конструктор перечисления
(Shape::Circle(...)); со строчной — квалифицированный вызов функции модуля (math::add(...)).
10. Модули
// math.glyph —— модуль
@module my.math
@pub @fn add(a: Int64, b: Int64) -> Int64 { return a + b; }
// main.glyph —— entry
@use my.math;
@fn main() -> Void {
let sum: Int64 = my.math::add(1, 2);
}
- Entry-файл (тот, что передаётся в
glyphc run/compile) не префиксуется. - Имена импортированных модулей префиксуются (
add→my_math_add); ссылки на собственные функции внутри модуля переписываются автоматически. @pub— публичность; по умолчанию символы приватные, но в v1.0 маркировка не является жёстким ограничением.
11. Тесты и ассерты
См. docs/testing.md.
12. Конкурентность
Вызов @fn async не выполняет тело, а возвращает ленивый хендл Async<T>.
spawn h; ставит задачу в общую FIFO-очередь пула воркеров (идемпотентно);
пул — это M:N (M задач на N потоках), N = число ядер, не больше 8,
задаётся через переменную окружения GLYPH_WORKERS=K. h await ждёт
результат: выполняет синхронно, если хендл ещё не запущен; иначе ждёт
завершения на condvar. Ждущий поток «помогает» пулу — пока цель не готова,
он разбирает очередь и исполняет другие задачи (work-sharing), поэтому
вложенные await внутри async-функций не блокируют пул намертво. Повторный
await возвращает кэшированное значение:
@fn async fetch(url: String) -> String {
return url;
}
@fn main() {
let h: Async<String> = fetch("http://x");
spawn h;
println(h await);
}
List/Map через границу воркера передаются по значению-структуре:
трамплин кладёт struct в ячейку, а каждый await делает её глубинную копию
(GlyphList/GlyphMap memcpy-данные + свежий refcount). Поэтому хендл
хранит приватный экземпляр результата: мутация и free() полученного по
await значения безопасны и портят только эту копию; повторный await снова
даёт исходное значение. (Ячейка результата хендла не освобождается —
тот же класс утечки, что и у Int64/String: рантайм без GC.)
Типизированные каналы создаются конструктором Channel<T>(capacity):
ёмкость 0 — rendezvous (прямая передача без буфера), > 0 — буфер.
send(v) блокирует при полном буфере и возвращает false, если канал закрыт;
recv() блокирует при пустом буфере и возвращает None, когда канал закрыт
и пуст; close() будит всех ожидающих:
@fn async produce(ch: Channel<Int64>) -> Int64 {
ch.send(1);
ch.send(2);
ch.close();
return 2;
}
@fn main() {
let ch: Channel<Int64> = Channel<Int64>(4);
spawn produce(ch);
let m: Option<Int64> = ch.recv();
let v: Int64 = match m {
| Some(x) => x,
| None => -1
};
print_int(v);
}
select: ожидание первого готового события
select { ... } ждёт одно из нескольких событий и выполняет соответствующий
блок. Рука — это либо | имя: Тип <- канал.recv(), либо
| имя: Тип <- хендл await, либо | timeout(мс), либо | default.
Срабатывает первый по порядку источник, у которого данные уже готовы;
для recv из буферизированного канала и для await заведение задачи в пул
происходит автоматически. timeout(мс) срабатывает, если за мс миллисекунд
ни одно событие не стало готовым; default — если прямо сейчас готовых нет.
timeout и default нельзя указывать одновременно (в каждом select),
timeout-рук тоже не может быть несколько. Логика ожидания — пулинг с шагом
~1 мс, поэтому таймауты честные, а пул воркеров продолжает работать:
@fn async worker(ch: Channel<Int64>) -> Int64 {
ch.send(42);
ch.close();
return 42;
}
@fn main() {
let ch: Channel<Int64> = Channel<Int64>(1);
let h: Async<Int64> = worker(ch);
select {
| v: Int64 <- ch.recv() => print_int(v),
| v: Int64 <- h await => print_int(v),
| timeout(100) => print_int(-1)
}
}
recv из закрытого и опустошённого канала — готов немедленно и возвращает
нулевое значение типа (аналог None).
Ограничения: рантайм без GC (хендлы, каналы и списки не освобождаются);
Option/Result-боксы: временные (результат вызова прямо в match)
освобождаются автоматически, именованные — через drop(box); рука select
должна быть именно recv() или await (произвольные выражения, включая
recv_timeout, не допускаются); потоки не вытесняются на уровне задачи —
долгий вычислительный корень задачи занимает воркера целиком (через вложенные
await воркер переключается на другие задачи); async-generic функции
запрещены; send(&ref) запрещён типом (адрес стека нельзя передавать
между потоками). Значения List/Map в канале передаются копией структуры с
разделяемым буфером (COW, как String): после send отправитель не должен
переиспользовать или освобождать значение — владелец приёмник.
13. Стандартная библиотека
Библиотека доступна без @use; сигнатуры объявлены в компиляторе,
тела — в stdlib/.
print(msg: String) -> Void
println(msg: String) -> Void
eprintln(msg: String) -> Void
print_int(value: Int64) -> Void
print_float(value: Float64) -> Void
print_bool(value: Bool) -> Void
read_line() -> String
sqrt(x: Float64) -> Float64
pow(base: Float64, exp: Float64) -> Float64
abs(x: Float64) -> Float64
floor(x: Float64) -> Int64
ceil(x: Float64) -> Int64
round(x: Float64) -> Int64
min(a: Float64, b: Float64) -> Float64
max(a: Float64, b: Float64) -> Float64
clamp(value: Float64, low: Float64, high: Float64) -> Float64
log(x: Float64) -> Float64 log2(x: Float64) -> Float64
log10(x: Float64) -> Float64 sin(x: Float64) -> Float64
cos(x: Float64) -> Float64 tan(x: Float64) -> Float64
len(s: String) -> Int64
substring(s: String, start: Int64, end: Int64) -> String
contains(s: String, sub: String) -> Bool
starts_with(s: String, prefix: String) -> Bool
ends_with(s: String, suffix: String) -> Bool
replace(s: String, from: String, to: String) -> String
split(s: String, delimiter: String) -> String
trim(s: String) -> String
to_upper(s: String) -> String to_lower(s: String) -> String
char_at(s: String, index: Int64) -> String
parse_int(s: String) -> Option<Int64>
parse_float(s: String) -> Option<Float64>
int_to_string(value: Int64) -> String
float_to_string(value: Float64) -> String
bool_to_string(value: Bool) -> String
to_int(s: String) -> Option<Int64>
to_float(s: String) -> Option<Float64>
to_bool(s: String) -> Option<Bool>
file_read(path: String) -> String
file_write(path: String, content: String) -> Bool
file_append(path: String, content: String) -> Bool
file_exists(path: String) -> Bool
create_dir(path: String) -> Bool
remove_file(path: String) -> Bool
list_dir(path: String) -> String
file_copy(src: String, dst: String) -> Bool
file_rename(old: String, new: String) -> Bool
alloc(size: Int64) -> Bytes
free(ptr: Bytes) -> Void
memcpy(dst: Bytes, src: Bytes, size: Int64) -> Void
memset(dst: Bytes, value: Int64, size: Int64) -> Void
assert(condition: Bool, msg: String) -> Void
assert_eq(a: Int64|UInt64|Float64|Bool|String, b: ..., msg: String) -> Void
assert_ne(a: ..., b: ..., msg: String) -> Void
assert_true(condition: Bool, msg: String) -> Void
assert_false(condition: Bool, msg: String) -> Void
NeuralScript (.ns) — Tensor DSL inside glyphc
NeuralScript is the .ns DSL for neural networks, integrated as glyphc nns.
It type-checks tensor shapes at compile time and emits a self-contained C++,
CPU SIMD-reference, or CUDA translation unit plus an optional C-ABI runtime
header.
Source files: examples/nns/*.ns (mlp, transformer, dense, static, moe, neumoe).
1. Quick start
glyphc nns examples/nns/mlp.ns --check # shape/type check only
glyphc nns examples/nns/mlp.ns --cpp # C++ reference to stdout
glyphc nns examples/nns/mlp.ns --cpp --runtime # C++ + ns_runtime.h header
glyphc nns examples/nns/mlp.ns --cuda --runtime # CUDA backend + header
glyphc nns examples/nns/mlp.ns --mlir # MLIR IR dump (debug)
glyphc nns examples/nns/mlp.ns --cuda --fp16 --runtime # experimental fp16 header/ABI
glyphc nns examples/nns/mlp.ns --cpp --runtime -o model.cpp # write output directly
# Compile the emitted driver with the generic host:
glyphc nns examples/nns/mlp.ns --cpp --runtime > /tmp/model.cpp
g++ /tmp/model.cpp examples/nns/host.cpp -I . -o /tmp/driver && /tmp/driver
All diagnostics use Static shape/type errors: with line:col locations;
codegen shape errors (e.g. dynamic dims not lowered) surface the same way.
2. Grammar
2.1 Type aliases and dimensions
type Batch = Dynamic // symbolic batch dim, resolved at call-site
type Features = 784 // constant dim
type Classes = 10
type Seq = Dynamic
type D = 512
type H = 8
// Shorthand forms:
type Hidden = 2048
DimExpr is one of:
Dynamic— unknown until runtime (batch / sequence length)Constant— integer literal (784,10,512)Symbolic— alias name that resolves to Dynamic or Constant viatypetable
Tensor type syntax (shape checker normalises alias names):
Tensor[Batch, Features] float32
Tensor[S, C] float32
Tensor[Seq, D] float32
Supported dtypes (src/nns/ast.rs:Dtype):
| Token | Dtype |
|---|---|
float16 | Parsed as Float16; codegen currently rejects native half kernels |
float32 | Float32 |
float64 | Float64 |
int8 / int16 / int32 / int64 | Int* |
fp8 / fp4 | Parsed as Fp8 / Fp4; codegen currently rejects these types |
bool | Bool |
Scalars use the same dtype tokens without Tensor[...].
2.2 Network, layers, forward
network MLPClassifier {
input: Tensor[Batch, Features] float32
output: Tensor[Batch, Classes] float32
layer fc1 = Dense(in: Features, out: 512, activation: ReLU)
layer drop = Dropout(rate: 0.1)
layer fc2 = Dense(in: 512, out: Classes, activation: Identity)
forward(x) {
return x -> fc1 -> drop -> fc2
}
}
Grammar elements:
network <Name> { ... }— one network per file is typical; multiple allowed.input:/output:— declared I/O tensor types; used to typeforwardparams and verify the return shape.layer <name> = <LayerType>(params...)— layer declarations.
Layer types and params (src/nns/shape_checker.rs:LayerRule):
| Layer | Params | Shape rule |
|---|---|---|
Dense / Linear | in:, out:/out_features:, activation: | [..., in] -> [..., out] |
Dropout | rate: | shape-preserving |
LayerNorm | (none / d_model:) | shape-preserving |
Attention / MultiHeadAttention | d_model:/dim:, num_heads:/heads:, causal: | shape-preserving ([B, S, D] -> [B, S, D]) |
Embedding | vocab_size:, d_model: | [..., S] -> [..., S, D] (appends dim) |
MoE / MixtureOfExperts | experts:/num_experts:, ffn_dim:, initial_experts: | shape-preserving |
activation: values are free-form strings; codegen handles ReLU, GELU,
Identity, Softmax, etc. Param values may be literals or symbolic alias names.
Forward:
forward(x) {
var h = x -> emb
var a = h -> attn
var n = a -> ln1
var g = n -> mlp1 -> mlp2 // pipeline chain desugars left-to-right
var s = g + n // elementwise add (same shape)
return s -> ln2 -> fc // return expression must match declared output rank+dims
}
->is the pipeline operator (PipelineOp):x -> fc1 -> dropapplies each layer’s shape rule sequentially; unknown stages are passthrough.+is elementwise tensor addition (rank+dim must unify).var/letbindings insideforwardare supported.forwardparams with no explicit type default to the network’sinputtype.
Larger example — 7-block transformer with residual adds
(examples/nns/dense.ns):
network DenseWide {
input: Tensor[S] float32
output: Tensor[S, C] float32
layer emb = Embedding(vocab_size: V, d_model: D)
layer b0_attn = Attention(d_model: D, num_heads: H, causal: true)
layer b0_f1 = Dense(in: D, out: FF, activation: GELU)
layer b0_f2 = Dense(in: FF, out: D, activation: Identity)
// ... b1..b6 identical blocks ...
layer head = Dense(in: D, out: C, activation: Identity)
forward(x) {
var h = x -> emb
var t0a = h -> b0_attn
var t0b = h + t0a
var t0c = t0b -> b0_f1
var t0d = t0c -> b0_f2
var t0 = t0b + t0d
// ... repeat ...
return t6 -> head
}
}
2.3 train { grad {} } — AOT autodiff
train(batch_x: Tensor[Batch, Features],
batch_y: Tensor[Batch, Classes]) -> float32 {
grad {
var hidden = batch_x @ fc1 // matmul: [B,784] @ [784,512] -> [B,512]
var act = relu(hidden) // elementwise activation
var preds = act @ fc2 // [B,512] @ [512,10] -> [B,10]
var loss = cross_entropy(preds, batch_y)
}
return loss
}
train declares the optimisation step. Inside grad { ... }:
@is matrix multiply (MatmulOp); shape checker enforces 2-D, inner dims unify.- Activations:
relu,gelu,sigmoid,tanh,silu,swish,softmax,leaky_relu,dropout,identity— shape-preserving. cross_entropy(preds, labels)expectsTensor[B, C]both; returns scalarfloat32.- Data-movement builtins:
slice,index,scatter,concat,transpose,reshapewith special shape inference (seeshape_checker.rs). - Layer names (
fc1,fc2) bind to their weight tensors[in, out]for autodiff. No duplicatew1/w2— the network owns one weight per layer.
The compiler lowers grad {} to a reverse-mode tape, fuses kernels
(mlir/fusion), then appends the optimizer. Scalars promote via the same
backend as inference. Return value of train must be the scalar loss.
In a full model the same forward body is usually replayed inside grad
(examples/nns/transformer.ns, dense.ns):
train(x: Tensor[S], labels: Tensor[S, C]) -> float32 {
grad {
var h = x -> emb
var a = h -> attn
var n1 = a -> ln1
var g = n1 -> mlp1 -> mlp2
var r = g + n1
var n2 = r -> ln2
var logits = n2 -> fc
var loss = cross_entropy(logits, labels)
}
return loss
}
3. Shape and type checking
The checker (src/nns/shape_checker.rs) runs in phases:
- Collect
typealiases into a dim table. - Register
layerinference rules. - Walk
network→forward/trainblocks, inferringTensorTypeperExpr.inferred_typeand unifying dims.
Unification rules:
Dynamicunifies with anything.Const == Constmust match value-wise, elseDimension mismatch.Symbolic == Symbolicwith different names merges symbolic constraints.Symbolic == Constbinds the symbol to that constant (seeunify_dim,merge_symbols,record_symbol_const).- Rank must match for elementwise ops and
cross_entropy;@requires both operands 2-D with equal inner dims.
The declared output rank is verified against forward return expression.
Failure emits:
Static shape/type errors:
4:12 Dimension mismatch: constant 784 vs 512
7:5 Network forward() produces rank-2 but declared output has rank-3
Only after a clean check does the compiler proceed to MLIR lowering; --check
stops here with Shape checking passed..
4. Backends and flags
| Flag | Effect |
|---|---|
--check | Only run shape/type checking, no codegen. |
--mlir | Lower to MLIR and dump IR to stdout (MLIRCompiler::dump). No C++ emit. |
--cpp | Emit CPU C++ reference backend (TargetBackend::CpuCxx). |
--simd | Select the CPU SIMD target; currently emits the same scalar reference with an explicit marker. |
--cuda | Emit CUDA backend (TargetBackend::Cuda). Default when neither --cpp nor --cuda is set; --cuda wins if both are set. |
--runtime | Prepend the C-ABI runtime header (ns_runtime.h, from src/nns/runtime/ns_runtime.rs) to the emitted source. Required when linking against host.cpp. Without it the output is just the graph kernels. |
--fp16 | Experimental CUDA fp16 header/ABI mode. The current kernels still use the float reference path; native half arithmetic and tuned half2 launches are not enabled yet. |
-o <file> / --output <file> | Write generated source to a file instead of stdout. |
The CLI stages are Lex -> Parse -> ShapeCheck -> MLIR -> Fusion -> Codegen.
Fusion (FusionPass::run) fuses matmul + activation / layernorm / bias
groups and reports Fusion: N groups fused on stderr for codegen runs.
Rocm and Metal backend variants are API-level placeholders that delegate to
the CUDA emitter with a backend marker; they are not native HIP or MSL backends
yet. The CLI currently exposes the stable CPU and CUDA targets.
Example — inspect all stages for mlp.ns:
glyphc nns examples/nns/mlp.ns --check # 1) types only
glyphc nns examples/nns/mlp.ns --mlir # 2) MLIR dump
glyphc nns examples/nns/mlp.ns --cpp # 3) kernels only (no header)
glyphc nns examples/nns/mlp.ns --cpp --runtime # 4) kernels + header (host-linkable)
5. --runtime C-ABI — symbol list
When --runtime is set the emitted file starts with ns_runtime.h and then
defines the following extern "C" interface (see src/nns/runtime/ns_runtime.rs
for the exact header text, version 1.2.0):
Weight layout
typedef struct ns_weight_desc {
const char* name; // e.g. "fc1_w", "emb_w", "moe_g_w"
size_t offset; // float offset in the blob
size_t count; // number of floats
} ns_weight_desc;
typedef struct ns_weight_layout {
size_t num_weights;
const ns_weight_desc* desc;
} ns_weight_layout;
Weights are one contiguous host-owned float blob [w0 | w1 | ... | wN-1]
in model-definition order.
Lifecycle / inference
ns_model* ns_runtime_init(const float* weights, size_t num_floats);
void ns_free(ns_model* m);
int ns_eval_infer(ns_model* m, const float* input, float* output,
size_t input_numel);
size_t ns_model_output_numel(const ns_model* m, size_t input_numel);
size_t ns_model_weight_count(const ns_model* m);
size_t ns_weight_count_static(void); // no model needed, compile-time constant
int ns_model_get_weights(const ns_model* m, float* out, size_t n);
const ns_weight_layout* ns_model_layout(const ns_model* m);
ns_weight_count_static()lets the host size its buffer before init.input_numelisbatch * in_cols(row-major, batch-leading).ns_model_output_numeltells the host how large the output buffer must be.
Checkpoints
int ns_save_checkpoint(const ns_model* m, const char* path);
int ns_load_checkpoint(ns_model* m, const char* path);
Format: NSM1 (weights only) or NSM2 (weights + MoE mask {n_layers, capacity, mask}).
NSM1 loads with all experts alive.
MoE expert lifecycle (only if graph has a MoE layer)
size_t ns_expert_count(const ns_model* m);
size_t ns_expert_birth(ns_model* m, int n); // activate n dead slots (perturbed copy)
size_t ns_expert_merge(ns_model* m, int a, int b); // average b into a, deactivate b
size_t ns_expert_kill(ns_model* m, int k); // deactivate k
Capacity (ns_moe_cap) is compile-time constant (gate row [D,cap] + per-expert
FFN weights); only liveness is runtime. Exactly one MoE layer per model.
AOT training (only if train is defined)
int ns_runtime_train_step(ns_model* m, const float* input, const float* labels,
size_t input_numel, float* loss_out, float lr);
int ns_objective_loss(ns_model* m, const float* input, const float* labels,
size_t input_numel, float* loss_out);
Training and inference share the same weight blob; ns_runtime_train_step
runs forward + backward + one optimizer step (Muon for matrices, AdamW
otherwise) and writes updated weights back into the model. ns_objective_loss
is forward-only.
6. Minimal host.cpp walkthrough
Reference driver: examples/nns/host.cpp (also examples/nns/neumoe/host.cpp
for the corpus/MoE study). It is generic over any compiled .ns — the weight
layout and MoE lifecycle are queried at runtime.
#include "ns/runtime/ns_runtime.h"
#include <vector>
#include <cstdio>
int main() {
// 1. Size the blob before creating a model.
const size_t nw = ns_weight_count_static();
std::printf("weight_count=%zu (%.2fM)\n", nw, nw / 1e6f);
std::vector<float> w(nw);
// 2. Need a layout to init correctly — build a temp model to fetch it.
// (The real host rolls Xavier / N(0,0.02) per desc.name — see host.cpp:init_weights)
ns_model* tmp = ns_runtime_init(w.data(), nw);
const ns_weight_layout* lay = ns_model_layout(tmp);
// init_weights(w, lay) — per-weight-name strategy:
// emb_w/head_w -> N(0,0.02), _attn_ -> N(0,0.02)*0.88,
// moe_g_w -> N(0,0.02), f1/f2 -> Xavier uniform
ns_free(tmp);
// 3. Create the real model with the rolled blob.
ns_model* m = ns_runtime_init(w.data(), nw);
std::printf("experts_initial=%zu\n", ns_expert_count(m));
// 4. Training loop (mlp-style: batch of features + one-hot labels).
// For the LM study (dense/neumoe) seq_len = L, vocab = 10240,
// xs is float-encoded token ids [L-1], ys is one-hot [L-1, V].
std::vector<float> xs(L - 1), ys((L - 1) * 10240), out((L - 1) * 10240);
for (int step = 0; step < steps; ++step) {
float loss = 0.f;
float lr = /* cosine with warmup, see host.cpp:lr_at */;
ns_runtime_train_step(m, xs.data(), ys.data(), L - 1, &loss, lr);
// Optional: grow MoE gradually
if (step % 50 == 0 && ns_expert_count(m) < 9)
ns_expert_birth(m, 1);
// Validation: forward-only path
float vloss = 0.f;
ns_objective_loss(m, xs.data(), ys.data(), L - 1, &vloss);
ns_eval_infer(m, xs.data(), out.data(), L - 1); // logits
}
// 5. Persist. NSM2 captures MoE mask.
ns_save_checkpoint(m, "model.nsm2");
ns_free(m);
}
Build (CPU):
glyphc nns examples/nns/mlp.ns --cpp --runtime > /tmp/model.cpp
g++ /tmp/model.cpp examples/nns/host.cpp -I . -o /tmp/driver && /tmp/driver data/ 1000 out/
Build (CUDA, as in examples/nns/neumoe/run.sh):
glyphc nns examples/nns/neumoe/neumoe.ns --cuda --runtime > /tmp/model.cu
nvcc /tmp/model.cu examples/nns/neumoe/host.cpp -I . -arch=sm_75 -o /tmp/driver
The same host.cpp drives NeuMoE (growing ns_expert_birth), Static-MoE
(K=9 from step 0), and Dense (no MoE) by probing ns_expert_count.
7. See also
src/nns/— lexer, parser,shape_checker.rs,mlir/,codegen/,runtime/ns_runtime.rsexamples/nns/mlp.ns— minimal classifier withtrain{grad{}}examples/nns/transformer.ns— embedding + attention + layernorm + MLPexamples/nns/dense.ns— 7-block transformer with wide final FFNdocs/language_en.md/docs/language.md— Glyph (.glyph) reference
Editor Setup
VS Code
There is no marketplace release yet, but the repository contains a minimal
VS Code extension with the grammar and LSP client in
editors/vscode-glyph. You can also use the
standalone settings below.
1. File association (settings.json)
{
"files.associations": {
"*.glyph": "glyph",
"*.ns": "neural-script"
},
"[glyph]": {
"editor.tabSize": 4,
"editor.insertSpaces": true,
"editor.wordWrap": "off"
}
}
2. TextMate grammar stub
Create .vscode/glyph.tmLanguage.json in your workspace (or inside a local
extension syntaxes/glyph.tmLanguage.json):
{
"scopeName": "source.glyph",
"name": "Glyph",
"fileTypes": ["glyph"],
"patterns": [
{ "include": "#comments" },
{ "include": "#attributes" },
{ "include": "#keywords" },
{ "include": "#types" },
{ "include": "#strings" },
{ "include": "#numbers" }
],
"repository": {
"comments": {
"patterns": [
{ "name": "comment.line.double-slash.glyph", "match": "//.*$" },
{ "name": "comment.block.glyph", "begin": "/\\*", "end": "\\*/" }
]
},
"attributes": {
"patterns": [
{
"name": "keyword.other.attribute.glyph",
"match": "@(module|use|fn|struct|enum|impl|const|pub|test)\\b"
},
{ "name": "keyword.other.guard.glyph", "match": "#guard\\b" }
]
},
"keywords": {
"patterns": [
{
"name": "keyword.control.glyph",
"match": "\\b(let|mut|if|else|match|while|loop|for|in|return|break|continue|spawn|await|select|timeout|default|as)\\b"
},
{
"name": "constant.language.glyph",
"match": "\\b(true|false|None|Some|Ok|Err)\\b"
}
]
},
"types": {
"patterns": [
{
"name": "entity.name.type.glyph",
"match": "\\b(Int64|UInt64|Float64|Bool|String|Bytes|List|Map|Option|Result|Channel|Async|Void)\\b"
}
]
},
"strings": {
"patterns": [
{ "name": "string.quoted.double.glyph", "begin": "\"", "end": "\"", "patterns": [{ "include": "#escapes" }] }
]
},
"numbers": {
"patterns": [
{ "name": "constant.numeric.glyph", "match": "\\b0x[0-9a-fA-F_]+|\\b\\d[\\d_]*\\.?\\d*[\\d_]*\\b" }
]
},
"escapes": {
"patterns": [{ "name": "constant.character.escape.glyph", "match": "\\\\." }]
}
}
}
Add a matching source.neural-script grammar for *.ns if desired
(keywords: network, layer, forward, train, grad, Tensor, Dynamic,
type, plus dtype tokens).
3. LSP — glyphc lsp (or glyphc --lsp)
The compiler exposes a stdio LSP server. Both forms are supported:
glyphc lsp
# backward-compatible alias:
glyphc --lsp
Capabilities: textDocumentSync: Full, completionProvider (triggers @ # . :),
hoverProvider, definitionProvider, diagnosticProvider. Diagnostics cover
lexer / parser / typechecker with full spans (line:col-col on type errors,
line:col on lex/parse).
VS Code — vscode-languageclient config
Example extension.js for a local extension (or use
vscode-languageclient + LanguageClient):
const { LanguageClient, TransportKind } = require('vscode-languageclient/node');
function activate(ctx) {
const serverOptions = {
command: 'glyphc',
args: ['lsp'],
transport: TransportKind.stdio
};
const clientOptions = {
documentSelector: [{ scheme: 'file', language: 'glyph' }],
synchronize: { fileEvents: [] }
};
const client = new LanguageClient('glyph', 'Glyph LSP', serverOptions, clientOptions);
ctx.subscriptions.push(client.start());
}
exports.activate = activate;
Alternatively, wire glyphc lsp as a generic LSP via extensions like
vscode-lsp-generic or lsp-bridge:
{
"lsp.servers": {
"glyph": {
"command": "glyphc",
"args": ["lsp"],
"filetypes": ["glyph"]
}
}
}
Check it works: open a .glyph file with a type error — diagnostics should
appear on save/change (full-sync).
Neovim (nvim-lspconfig)
glyphc lsp is a plain stdio server. Register it as a custom lspconfig
server. Requires neovim/nvim-lspconfig.
Minimal init.lua
local lspconfig = require('lspconfig')
local configs = require('lspconfig.configs')
if not configs.glyph then
configs.glyph = {
default_config = {
cmd = { 'glyphc', 'lsp' },
filetypes = { 'glyph' },
root_dir = function(fname)
return lspconfig.util.root_pattern('glyph.toml', '.git')(fname)
or lspconfig.util.path.dirname(fname)
end,
single_file_support = true,
settings = {},
},
}
end
lspconfig.glyph.setup {
on_attach = function(client, bufnr)
-- optional: hover / goto-def keymaps
vim.keymap.set('n', 'K', vim.lsp.buf.hover, { buffer = bufnr })
vim.keymap.set('n', 'gd', vim.lsp.buf.definition, { buffer = bufnr })
vim.keymap.set('n', '<C-Space>', function() vim.lsp.buf.completion() end,
{ buffer = bufnr, mode = 'i' })
end,
}
-- .glyph filetype
vim.filetype.add({ extension = { glyph = 'glyph', ns = 'neural_script' } })
With lazy.nvim:
{
'neovim/nvim-lspconfig',
config = function()
-- same configs.glyph block as above
end,
}
Backward-compatible flag
Older configurations may still use glyphc --lsp; it is an alias for
glyphc lsp. New configurations should prefer the subcommand form.
What you get
- Live diagnostics: lexer errors (unterminated string, bad number), parser
errors (
Unexpected token), type errors with full spans (line:col-col: Undefined variable: nope). Seesrc/lsp/mod.rs:analyze. - Hover (
textDocument/hover): keyword docs forselect,spawn,await,Channel,Async,Result,Option, etc. - Completion (
textDocument/completion):@module,@fn,@struct,@enum,let,match,spawn/await/select/timeout/default, … - Go-to-definition (
textDocument/definition): locallet/ param → definition in the same item, else globalfn/structname.
Limitations: diagnostics are per-file (no cross-module workspace analysis yet); text sync is Full (whole buffer on each change).
Troubleshooting
glyphc: command not found— build first:cargo build --release→target/release/glyphconPATH.- No diagnostics: ensure
filetypesincludesglyphand the server is attached (:LspInfoin Neovim, Output panel in VS Code). - Hover empty: cursor must be on a word (
word_atusesis_alphanumeric || _ || # || @).
Тестирование в Glyph
Атрибут @test + встроенные ассерты + подкоманда glyphc test.
Атрибут @test
Любая функция может быть тестом:
@test @fn test_add() -> Void {
// ...
}
Правила:
- функция не принимает параметров;
- возвращает
Void; - грамматика
@test @fn— любой порядок с@pub(@pub @test @fn,@test @pub @fn).
Нарушение проверяется на этапе типизации:
Type error: Test function 'test_with_param' must not have parameters
Type error: Test function 'test_bad_ret' must return Void (found Int64)
Ассерты
Доступны без @use (аналоги в stdlib/assert.glyph):
assert(condition: Bool, msg: String) -> Void
assert_true(condition: Bool, msg: String) -> Void
assert_false(condition: Bool, msg: String) -> Void
assert_eq(a, b, msg: String) -> Void // a == b
assert_ne(a, b, msg: String) -> Void // a != b
assert_eq/assert_ne полиморфны по типам: оба аргумента должны быть одного
типа из Int64, UInt64, Float64, Bool, String. При провале выводится msg.
В обычной (не тестовой) сборке проваленный ассерт аварийно завершает программу.
glyphc test
Как это работает:
glyphc test # сканирует ./src рекурсивно, ищет *.glyph с @test
glyphc test -i src/calc.glyph # только один файл
glyphc test --compiler clang # сборщик
glyphc test --opt -O0 # уровень оптимизации
Для каждого файла:
- ищутся все
@test-функции (файлы без тестов пропускаются); - файл компилируется в обычном режиме + генерируется специальный C-рантайм,
который выполняет каждый тест в изоляции (
setjmp/longjmp) — падение одного теста не прерывает остальные; - построенная программа возвращает протокол в stdout.
Формат протокола от glyphc test:
Running 8 test(s) in src/calc.glyph...
[ok] test_add
[ok] test_add_negative
[FAILED] test_fail
this test should fail
Result: FAILED (2 file(s), 9 passed, 1 failed)
- exit code
0— все тесты прошли; - exit code
1— есть упавшие тесты (берётся программа, собравшаяся успешно); - ошибки компиляции/типов по конкретному файлу не прерывают остальные файлы, результат суммируется.
Пример проекта
examples/tests/ — законченный тестовый проект из двух модулей:
examples/tests/
glyph.toml
src/calc.glyph @module calc; 8 тестов (7 успешных + 1 намеренно падающий)
src/main.glyph @use calc; 2 теста, проверяют вызовы через calc::add
Запуск:
cd examples/tests
glyphc test
# Result: FAILED (2 file(s), 9 passed, 1 failed) — test_fail упал намеренно
glyphc test -i src/main.glyph
# Result: OK (1 file(s), 2 passed, 0 failed)
Регрессия компилятора
Юнит-тесты самого компилятора:
cargo test
Полный прогон всех примеров языка:
examples/run_all.sh
Glyph vs Zig / Odin / Mojo / tinygrad — Honest Comparison
Scope: where Glyph actually sits among its closest neighbours. Not a benchmark shootout and not FUD — each tool optimises for a different trade-off. Read the per-system notes before picking.
Summary table
| Dimension | Glyph | Zig | Odin | Mojo | tinygrad |
|---|---|---|---|---|---|
| Primary goal | Small, C-transpiled application language with an integrated tensor DSL (.ns) and M:N concurrency | Systems language, explicit control, no hidden allocations | Pragmatic systems language (Jai heritage), batteries-included | Python-superset for AI, CPU/GPU kernels, Python interop | Minimal deep-learning framework (Python), <10k LOC |
| Runtime model | No GC, refcounted List/Map, drop(box) for Option/Result; async handles/channels not auto-freed | Manual memory, no GC, explicit allocators | Manual + optional allocators, no GC | Borrowed Python runtime; Mojo structs can be __del__-managed | Python runtime, lazy UOp graph |
| Compilation | Transpiles to C (GNU statement exprs) → gcc/clang (-O2, -pthread) | LLVM via own toolchain, self-hosted | LLVM | MLIR → LLVM / custom GPU codegen | Python JIT → LLVM / METAL / CUDA / CL / HIP |
| Concurrency | M:N worker pool, Async<T>, spawn/await, Channel<T>, select { timeout/default } | std.Thread, atomics, no built-in async | core:thread, core:sync | Mojo async (evolving), Python asyncio interop | Single-threaded Python loop + device queues |
| Type system | Static, monomorphized generics on functions only, Result/Option as boxed unions | comptime generics, error unions, optionals | Parametric polymorphism, Maybe / Error | Strong + grad. typing, traits, ownership | Python dynamic + shape tracking |
| Tensor / ML story | First-class: .ns (Tensor[Dims], network/layer/forward/train{grad{}}, shape checker, fused MLIR → C++/CUDA, C-ABI with MoE + checkpoint) | Libraries only (no DSL) | Libraries only | First-class: Tensor, SIMD, autotune, MAX Engine, GPU | First-class: tinygrad is the framework |
| Tooling | glyphc check/tokens/ast/run/build/test/fmt/nns --check/--cpp/--cuda/--runtime/--fp16 + glyphc lsp (hover, completion, diagnostics) | zig build, LSP (zls), zig fmt | odin build, odlsp | mojo build/run, mojo LSP, pixi | python -m pytest, TINY_BACKEND env |
| C interop | Generated C is plain gnu11; host links directly against ns_runtime.h | @cImport, @extern | foreign import | external_call, C FFI | ctypes / cffi via Python |
| Maturity | v2.1, single glyphc binary, 100 unit + 25 integration tests + examples/run_all.sh; former nsc merged in | Production — self-hosted compiler, many users | Production (used at Janga, FMOD) | Public preview → GA track; breaking changes still happen | Production for hobby/research, used in prod by some |
| Best fit | You want one binary that compiles both application logic and a parameterised model with compile-time shape checks and a C/CUDA host loop | You want a C replacement with comptime and tight control | You want a concise C replacement with pragmatic stdlib | You want Python-compatible AI code that lowers to fast kernels | You want to read/hack the whole DL stack in Python |
Per-system notes
Glyph
Strengths: single glyphc binary covers app + model; .ns catches rank/dim
errors before codegen (symbolic + Dynamic unify, MoE liveness at runtime);
emitted kernels are plain C++/CUDA that link against a minimal ns_runtime.h
(C-ABI version 1.2.0) — see docs/nns.md and examples/nns/host.cpp.
Trade-offs: runtime is unmanaged (no GC) — xs.free() and drop(box) are
manual and the use-after-free detector is statement-flow only (aliases inside
match etc. not tracked). Generics only on functions (no generic structs/enums,
no trait bounds; f::<T> not supported). Map iterates in hash order, keys are
String only. select polls at ~1 ms, async-generic fn unsupported.
LSP is Full-sync, per-file diagnostics.
Zig
Strengths: strongest story for explicit, auditable systems code; comptime is
more expressive than Glyph’s monomorphization; error handling via error unions
is ergonomic; allocator-aware stdlib.
Trade-offs: no built-in async or channels; no tensor DSL — all ML is via external libs. Build system and language churn faster than C. No GC either, but memory errors are caught differently (no statement-flow alias tracking for collections — you just use allocators).
When to prefer Zig: you need resource-constrained, allocation-transparent systems code and do not need in-language ML.
Odin
Strengths: very pragmatic, small language surface, good C FFI, concise. Comparable C-replacement niche to Zig but with different ergonomics.
Trade-offs: similar to Zig — no built-in tensor/ML, no M:N async, no select
multiplexer. Generics are limited functors/procedures.
When to prefer Odin: you value brevity and a stdlib that feels like a game engine toolkit.
Mojo
Strengths: Python syntax + systems speed; owns the full AI stack (SIMD,
autotune, GPU) with Python interop. For pure ML workloads it has a wider
kernel library than Glyph’s .ns.
Trade-offs: larger runtime/toolchain, faster-moving spec, ecosystem still
stabilising. Module/trait system is more general than Glyph’s @module/@use
prefix scheme but heavier for tiny programs.
When to prefer Mojo: the workload is AI-centric and you need Python interop or MAX Engine. Glyph wins when the program is a small native binary with a parameterised model and minimal dependencies.
tinygrad
Strengths: smallest DL stack you can read end-to-end; excellent for research, custom backends, and hacking the UOp graph. Python-native.
Trade-offs: it is a library, not a compiled language — performance and
deployment depend on Python + the chosen backend. No Glyph-style static shape
checker or generated C++ host; no glyphc fmt/test/lsp.
When to prefer tinygrad: you want to own the framework, in Python.
Guidance
- App + embedded model, one binary, static shapes, C/CUDA host → Glyph.
- Systems code with comptime and no hidden costs → Zig.
- Game/tools systems code, pragmatic stdlib → Odin.
- Python-compatible AI with kernels + interop → Mojo.
- Minimal Python DL framework to read/hack → tinygrad.
No system here is a superset of another; pick the one whose primary goal
matches yours. Contributions that improve the glyphc --lsp, .ns fusion
passes, and the Map/List runtime are especially welcome — see
docs/language_en.md and docs/nns.md.