Small, fast programming language where indexes are valid and values can't be shared.
In sloe you can refer to items and slices stored in consecutive memory in a safe and infallible way. With this you can for example represent tree-like data structures without segmented memory or plain indexes (which would need to check bounds and generations for safety).
fn Greet
.name name str
.buf buf Buf _origin, char
:
.buf Buf _origin, char
.span Span _origin
=
? Buf-add-str-chars .buf buf .new "Hello, " str [string]
? Buf-span-add-str-chars .. string .new name [string]
Buf-span-add .. string .new "!" char
↑ a Greet function which takes a name and a buffer to append the greeting to.
It appends the name and other strings to form and return a message span (range within the buffer).
explore examples in an online editor skip to more examples skip to syntax overview or look into the example-/ directories in this repo.
Install with (requires having rust installed)
cargo install --git https://codeberg.org/lue-bird/sloe sloePassing a value as an argument? Consumes it. Matching a value? Consumes it. Even variables holding plain numbers for example have to be explicitly duplicated when you need to use them in multiple places.
- values know when they aren't used anymore at compile time. Their memory is always explicitly reclaimed. No need for garbage collection or similar. Additionally, clean-up can be flexible, like a range of indexes freeing their memory by passing the containing collection
- values can be mutated internally without mutation being detectable
- guaranteeing properties like non-overlapping pointed memory regions can enable more optimizations, e.g. through llvm's
noalias - threads can only be joined once for example
It can sadly also feel clunky. Think e.g. Span-length which takes a span and gives back its size and the given span. Span-length could also return a changed Span behind your back. This flexibility can be an advantage but more importantly it sadly complicates tracking where a value changed.
An immutable view (like &Span in rust) would not have these difficulties.
The big advantage of this rule is how easy it is to understand and how much simpler and faster it is to statically analyze compared to lifetimes or similar.
Further reading if interested: "linear types"
A collection which can mark its indexes as unset without moving existing items around (thus invalidating their indexes).
This can be used to "return" memory which has become outdated or useless, for example with Buf-remove, Buf-span-rid for future reuse with for example Buf-insert.
(This functionality is optional. You can use a Buf for builders etc. which never try to reuse unset indexes before they are scrapped.)
Further reading if interested: "memory-reusing slot map"
Similar to allocators, you almost never access, alter or iterate their contained values directly.
Collections are seen as storage into which you can add items, build slices etc.
Whenever you do so, you'll get Slots and Spans that assert your permission to access and alter the referenced items as well as your responsibility to announce their release at some point.
Further reading if interested: "storage is not responsibility"
Every created collection has a unique origin. A value whose type contains an origin can't escape the scope of it's origin. This is checked at compile-time for the expression following origin creation but you'll likely realize it before then:
fn Some-buf . : Buf ??origin cannot even be annotated??, u32 =
^ buf-origin
? Buf-empty{u32} buf-origin [buf]
? Buf-add .buf buf .new 123 u32 [.buf buf .slot slot]
...
buf
# compiles
fn Add-some-values buf Buf _origin, u32 : Buf _origin, u32 =
? Buf-add .buf buf .new 123 u32 [.buf buf .slot slot]
...
buf
Further reading if interested: "origin"
^ some-origin-name creates a new variable of type Origin and a unique local type that's only valid in the current scope.
Like every other sloe value, an origin type can only be used once, so only for one collection.
# use a temporary collection contained within a scope
fn Use-buf . : u32 =
^ buf-origin
# create a buffer that can hold u32 items and give it the name buf
? Buf-empty{u32} buf-origin [buf]
# insert 123, destructure the resulting record
? Buf-add .buf buf .new 123 u32 [.buf buf .slot first-slot]
# without new slot, the referenced item is ours to modify or pop
? Buf-remove .buf buf .slot first-slot [.buf buf .item first]
# Consecutive slots connect into a span
? Buf-add-array .buf buf .new ; 456 u32 ; 789 u32 [.buf buf .span after-first]
...
first # = 123 u32
# different branches, different scopes
fn Use-opt opt Opt u32 : ... =
# this won't compile as their origins come from different branches
? (
? opt
['no .]
^ buf-origin
Buf-empty{u32} buf-origin
['yes number] (
^ buf-origin
? Buf-one .origin buf-origin .item number [.buf buf .slot slot]
...
buf
)
)
[buf]
# this will compile:
^ buf-origin
? (
? opt
['no .]
Buf-empty{u32} buf-origin
['yes number] (
? Buf-one .origin buf-origin .item number [.buf buf .slot slot]
...
buf
)
)
[buf]
...
# tree structure. every slot and span exclusively belongs to that expression.
# If passing so many origins seems annoying to you, check the documentation of Origin
ty Expression _expressions-origin, _patterns-origin, _chars-origin
'int i32
'string Opt Span _chars-origin
'buf Opt Span _expressions-origin
'call
.function Slot _expressions-origin
.arguments Span _expressions-origin
'lambda
.parameters Span _patterns-origin
.result Slot _expressions-origin
ty State _expressions-origin
# ...patterns, chars, positions etc
.expressions Buf _expressions-origin, Expression _expressions-origin
.root-expression Expression _expressions-origin
fn Initial-state
.expressions-origin expressions-origin Origin _expressions-origin, _expressions_part
: State (Origin _expressions-origin, _expressions_part) =
.expressions Buf-empty{Expression _expressions-origin} expressions-origin
.root-expression (..do parsing..)
fn State-to-interfaces-into
.interfaces interfaces Buf _interfaces-origin, Interface State _expressions-origin
.state state State _expressions-origin
: Buf _interfaces-origin, Interface State _expressions-origin =
? (
Buf-one
.origin interfaces-origin
.item 'console-log{Interface State _expressions-origin} "hello" str
)
[.slot slot .buf interfaces]
...
interfaces
(No need to understand the details at the end, it's just a small showcase to get a vague feel)
In case a function cannot scrap values like buffers at the end of its scope, we can pass origins or values referencing origins in:
fn Buf-empty{_item} Origin _origin, _part : Buf (Origin _origin, _part), _item
fn Buf-add .buf Buf _origin, _item .new _item : ...
Most initializer functions will return new collections from nothing, e.g. for persistent application state.
For most other functions, it's more common to pass in an existing collection that you want to edit (often also including a specific span).
(If you're wondering what _part is here: It enables creating an origin inside the function and still passing collections etc using that origin out of the function via Origin-erased. Look it up if you think the existing origin stuff is too restrictive)
explore more examples in an online editor or look at the example-/ directories in this repo for more real-world-like usage.
# line comment
# number (available types: p32, u32, i32, f32).
# Specifying a type is required
3.2 f32
# text of type str
"hello" str
# unicode scalar of type char
"a" char
# most identifiers
variable-or-field-or-variant-or-type-without-parameters-2012
# name of a function or type with parameters
Constructor-name
# function call. always 1 argument; no parens needed
Some-function argument
# very rarely functions may require type arguments {in braces}.
# Function declarations will explicitly list those {_parameters} (see later)
Some-function{type}{arguments} Inner-call-as-the-argument inner-call-argument
# record. You may know it as "struct".
# If field values themselves end in records they need to be parenthesized.
# The last field value can end in a record without needing to be parenthesized
.first-field first-value .second-field second-value
# empty record, like void/unit.
# commonly used for variants without a value,
# for lazily constructing a value, for empty state/context
# or as the result of functions like U32-rid
.
# ..spread a record into other fields
.field-1st value-1st .. one-existing-record .. another .field-2nd value-2nd
# temporary array. Rarely used
; first-item ; second-item ; third-item
# local function of type fn.
# the pattern must add a type to all variables
# can **not** use variables from the outer scope.
[parameter-pattern] result
# pattern variable
# appending a type is only necessary and allowed in function parameters
some-variable some-type
# pattern match, checked for exhaustiveness. expressions must be parenthesized if they themselves end in a query.
# The last case result does not need to be parenthesized
? value [first-case-pattern] first-result [second-case-pattern] second-result
# introduce a new origin (describes which collection slots and spans point into).
# The given name can be used as a variable and its unique local type.
# Below will create a variable `new-origin-name` of type `Origin new-origin-name, .`
^ new-origin-name expression-that uses new-origin-name
# introduce multiple new origins with the same unique local type but different part names.
# Below will create a variable view-origin of type
# .json Origin view-origin, .json .
# .html Origin view-origin, .html .
# .char Origin view-origin, .char .
# which you can then query with ? to get origin variables for the different fields.
# Not only can this reduce the amount of type variables floating about,
# it's also important for wrapping values into an `Origin-erased`
^ .json .html .char view-origin expression-that uses them
# project function declaration.
# For type variables in the result that aren't used in the input,
# functions require appended type parameters: {...}
fn Function-name{_potential}{_type-arguments}{_only-used-in-the-result}
parameter-pattern-with-types
: result-type
# optional documentation
# comment
=
result-expression
# type name without arguments. lowercase
u32
# type with multiple arguments. Uppercase name.
Buf origin, item
# arguments before the last must be parenthesized if they end in a type with arguments
My-function-type (Inner env), input, output
# declare a shorthand for an existing type
ty point .x i32 .y i32
# can also accept parameters
ty Pair _first, _second-parameter ..type using the type variables..
# a "choice type" that can come in different shapes ("variants")
# which each have a unique name and one associated value.
'first-option .
'second-option Buf _potential, u32
'third-option Type-name-alias _potential, _type-parameters
# creating a variant.
# The type in curlies can be a type alias or a choice type directly {'... ...}.
# The type does not have to include a variant with the currently constructed name
'some-variant-name{a-choice-type} its value
# variant pattern
'some-variant its value
Goal: coherent, practical and compact, avoiding parens and indentation especially for trailing syntax. Sloe is a wordy and explicit language, so any extra verbosity is not tolerable.
- clone this repo
- open the editor command panel
- "zed: install dev extension" or "gram: install extension from folder" and select the cloned-path-sloe/zed directory
Optionally for more precise syntax highlighting, add the setting "semantic_tokens": "combined" or "languages": { "sloe": { "semantic_tokens": "full" } }.
Optionally for a file icon in the project panel, open the editor command panel, select Icon theme selector: toggle and choose "sloe icon dark"/"sloe icon light".
- download https://codeberg.org/lue-bird/sloe/src/branch/main/vscode/sloe-0.1.0.vsix
- open the command bar at the top and select:
>Extensions: Install from VSIX
- clone this repo
- open
vscode/ - run
npm run packageto create the.vsix - open the command bar at the top and select:
>Extensions: Install from VSIX
write to ~/.config/helix/languages.toml:
[language-server.sloe]
command = "sloe lsp"
[[language]]
name = "sloe"
scope = "source.sloe"
injection-regex = "sloe"
file-types = ["sloe"]
indent = { tab-width = 4, unit = " " }
language-servers = [ "sloe" ]
auto-format = trueFor other editors, there's usually a way to specify sloe as the language server and or point to the directory tree-sitter/ in this repository.
As a user of sloe you can stop reading here. The rest is for developers and those interested in language design
nice short explainer in the austral language docs, article "must move types", "mutable value semantics".
Sloe once allowed values to be ignored ("leaked"/forgotten) making them "affine types", like rust owned values. This was changed as it was too easy to for example accidentally forget to handle a value in one query case but not the others. Better be safe and explicit.
Regarding the theoretical noalias optimization: languages sloe compiles to don't always exploit this fact.
In rust, a prominent example is slab. Comparison of various kinds of similar rust collections. There are even fast general purpose allocators based on this concept, for example zig's SmpAllocator or the rust crate "smmalloc".
There are many variations of this core idea (generational, external unset tracking, intrinsic free list, separate list of indexes, ...). Sloe does not have most problems these variations solve, so you can just think of sloe's Buf as a fast Vec<Option<Item>>.
The alternative to this would be to make tiny allocations for every slot and small span and to allow recursive types. This is convenient and not uncommon in languages like rust. However, sloe's goal is to do better here and to not group together storage and ownership over its items. Instead, we store a big array buffer of each kind and point into it.
I've heard this kind of decoupling being called "call-site dependency injection" which also perfectly applies to the idea of passing allocator, interner, concurrency runtime etc. around.
I really like this idea but understand that it cannot be implemented nicely in e.g. rust which needs to for example store its allocator in its value body to guarantee its content isn't scattered across different inaccessible allocator memories (and to satisfy Drop and to keep most of the existing function interfaces as well as convenience). Sloe solves this dilemma by assigning this unique origin at the high cost of user convenience.
In my opinion this isn't quite a solved problem and if you have other ideas, I warmly encourage you to explore and share them.
The insight "marking origin-specific types specific to code unique paths" has been hinted at in "The Unreasonable Effectiveness of Naming Integers". In sloe's case the unique origin types only exist at compile-time and can thus mark spans, slots, unset spans, unset slots, bufs etc. generically. Additionally it is checked that actually only one collection and its indexes are marked that way.
The idea of "fresh, distinct type instances by code" seems to generally be called "path-dependent types". In rust I know of 4 crates that implement this: compact_arena (safe, pragmatic, simple but bare-bones), indexing (safe, cumbersome, complicated) and generativity/typetoken explained in "the generativity pattern in rust" (general-purpose but relies on lifetimes). I've also seen ghostcell mentioned a lot in this context but have not investigated much.
The same idea but with runtime checking instead of compile-time checking can quite easily be implemented by storing an ID in each collection and the same id in each contained slot, and incrementing a global atomic variable (or similar) for the next available ID: example (apart from security this is hardly ever worth it for regular users, considering it is also slower).
to re-compile
cargo install --offline --debug --path . sloeWhat I'm unhappy with in the current design.
Writing these down has already helped a lot in coming up with fixes (e.g. Buf-span-add-own-span, Origin-erased etc. did not exist at one point but were created in response to now deleted list items).
And even if I'm unable to fix them, other people/teams might (in other projects)!
-
it seems quite natural to represent a span of structs as e.g.
.field-names Span _field-names .field-values Span _values. This pattern is more memory efficient and can reduce the amount of origins and Bufs necessary. The biggest missing convenience to make this attractive might be helpers to fold over many spans simultaneously. Honestly, the current "fold over one span and step through the rest withSpan-start" is annoying. I do not particularly like it as there is always "overspill" that needs to be handled. Additionally, this wastes memory for the duplicated memory (length u32 per extra Span) and wastes computation for unnecessarily handlingZig "fixes" this by both
- introducing special syntax and crashing at runtime if lengths differ
- only storing the length in one of multiple slices and documenting the expected length for the other start pointers
I think something like introducing
Span2 FirstOrigin, SecondOriginfor 2 up to maybe 5 makes sense. You'd be able to fold, access etc. them together and even split those up into separate spans whenever desired (but not join them back!). The sad thing is that this is positional and individual span slots then do not have an associated name. Also, how would this work with existing buf APIs? Something likeBuf2-opt-span-add -
minor: sometimes, you really own all the items of a buf in one place (especially when the buf items can be trivially copied). Splitting it into
Opt Span+Bufis annoying and wastes a bit of space (length is carried twice and start is always 0) -
by default, most passed arguments are quite fat on the stack (e.g.
Bufis 3-4.5 usize-wide and you may pass a bunch of them). Pointers are much thinner. This can in some parts be optimized by the target language compiler -
currently syntax is not full-word-search friendly. Think
_type-variableandminus-dash-hyphen -
the language is very sequential by design which disqualifies it from running fast on much of parallel computing e.g. GPUs, threads that share memory etc. Sloe is most likely not the right vehicle to explore this space, still it seems like a warning sign for a supposed "general-purpose language"
-
number types, vector/array types etc. are very underbaked in sloe. I need more real-world experience for their uses. Granted, sloe support for them is only realistic if rust (and zig) improve their support as well
-
even simple things often require a slog of things to type. Also, values, especially nested ones are "sticky" and getting rid of them is annoying. That sucks the fun out of programming, it's exhausting and slows you down for no apparent reason. Having a bunch of same-looking code also makes it harder to spot actually interesting or bugged parts, cancelling out many of the supposed benefits of linear types :(
I'm not blind to this feeling! It's the reason you might call sloe "great on paper, died to the real world". Do you have ideas of how this could be fixed in parts?
-
Currently Origin parts can only go 1 level deep. This means often you still do have to take separate type parameters for each part origin and origin erasing/unerasing more or less requires that all part-origin-dependent stuff gets covered.
How could sloe enable nested origins?
I've tried around a bit but can't seem to make this possible nicely:-- second try below -- For example to create a variable
inner-originof typeOrigin In (In unique-origin, .outer), .inner .,^ unique-origin {.outer .inner .} ? unique-origin [.outer .inner inner-origin]where
Inis a new core type that has no values and just exists to make types prettier. (Notice also how the Origin type itself only has 1 type parmeter now and types like Buf don' write out Origin everytime. Just In which reads much nicer)When unerasing, the exact type in origin can be taken to mean the parent:
fn Origin-unerase .erased Origin-erased _erased .unerase Fn (... .uneraser Origin-uneraser _origin), ... .uneraser Origin-uneraser _origin : ... fn Slot-origin-isolate Slot In _origin, _part : Origin-isolated _origin (Slot In erased, _part) fn Slot-origin-unerase .uneraser Origin-uneraser _origin .slot Slot (In erased _part) : ... .slot Slot In _origin, _partThis means though that
- we only unerase and erase 1 level deep (??). This must be addressed
- (less critical) This still means Origin/Slot/... by default need an
In unique, .argument
-- first try below -- For example to create a variable
inner-originof typeOrigin unique-origin, In (In ., .outer), .inner .,^ unique-origin {.outer .inner .} ? unique-origin [.outer .inner inner-origin]where
Inis a new core type that has no values and just exists to make types prettier.Question: Should simple origin creation also produce
Origin unique, In . .instead ofOrigin unique, .? The only reason I could see is that it would allowOrigin-uneraseto preserve the outer part of the given origin as the outer part of the unerased values. How would this work for sub-origins when isolating then?
-
IDE type and type diff error displays suck ass, mostly due to indentation being stripped. But markdown support seems to still be ways off for most editors for some reason. Anyone know a solution?
-
add
b8type to represent a byte. AddI32-2s-complement/U32/F32-to-b8s-decreasing/increasing-significanceand the other way around. AddB8-xor,B8-complement,B8-and,B8-or,B8-reinterpret-as-U32 -
add hashing primitives (or provide the means for users to implement them). Would for example be nice if index maps could in some way make use of Bufs, for example if it's like
ty Map _slots, _in-hash-order, _item .slots Buf _slots, Slot _in-hash-order .items Buf _in-hash-order, _item -
when in query case pattern record, suggest field name in completion
-
add field and variant rename and references
-
add code action for spreading a pattern variable
-
similarly, add "add remaining query cases" code action
-
suggest full parameter field patterns of existing project fns (just as rust does). This is super convenient, especially because stuff like
expressions Buf _expressions, Expression _expressions _patterns _typesdoesn't exactly roll easily over one's keyboard -
move collecting records and choice types to codegen phase, so that e.g. they are not unnecessarily collected for js or zig backend. Counter-argument: Most languages do not have a concept of structural composite types and repeating the job of collecting records for each one individually is error-prone and more work
-
add
Set _origin, _itemalong with add something likeMap _origin, _key, _value(or justMap _origin, _itemwhere key is derived from item) which still gives outSlot Origins for each entry but can be queried by key or similar.Map-emptywill require providing an.order (Fn .a _key .b _key, .a _key .b _key .order order) .dup (Fn _key, .a _key .b _key)or similar. Alternatively, check if implementing in userland via e.g. index map, AVL or red-black tree backed by a regularBufis fast enough -
consider adding
Buf-countingandslotwhich can reference a slot that is already in use:fn Buf-counting-slot-dup .buf Buf-counting _origin, _item .slot Slot _origin : .buf Buf-counting _origin, _item .a Slot _origin .b Slot _origin = # what about spans?In theory, this would enable graph structures, child-parent relations, doubly-linked lists, inlined string storage (although that would need e.g.
Set-counting) etc. Things I dislike with this design:- access via
Buf-counting-unsetwhich does not guarantee seems maybe too difficult (first un-occupy all known slots and even then there is no guarantee).Buf-counting-updateshould work nicely, especially for copiable types at the cost of: cannot access the buf at the same time and spooky action at a distance - every item is reference-counted. A slot to a known single-reference item cannot be represented. This is not a biggie because I don't know if there is a use for this
- maybe also -counting versions of map/set etc.
The alternative is of course to do
Slot-weakand generational indexes. However, this is un-usable for e.g. inlined string storage and also comes with overhead and even less guarantees.Open question of representation:
Buf<{ count: u32/16, item: Item }>: Finding unset slots takes linear time. Generally fast. Takes the least space on the stack{ items: Buf<Item>, counts: Buf<u32/16> }: Finding unset slots takes linear time. Generally fastest. A bit more error-prone than single buf{ items: Buf<Item>, unset: Buf<u32>, occoupied_counts: Buf<NonZeroU32/16> }: Finding unset slots takes constant time but doesn't feel deterministic. Generally fastest but unsetting is more expensive. More error-prone than single or double-buf. Takes most space on the stack{ items: Buf<Item>, counts: Buf<{ range: Range, count: u32/16 }>: Tough to handle and error-prone. Inefficient for cases where slots are handled one by one (no spans exist). Efficient for things like inline storage where spans are clearly defined.- the above but with unset ranges and occupied counts split
I think I prefer not storing counts in ranges, as for example for string interning, you could store counted spans in separate collections:
chars # of type Buf _chars, char names # of type Buf-counting _names, Span _chars ..other bufs pointing into chars, e.g. for number literals..This is likely the better option anyway (even though it "hops twice") as it makes searching for the right span possible (and reasonably fast)
- access via
-
introduce
ascii(in rust backed bystd::ascii::Asciwhich is currently experimental, in zig backed byu7), require char literals to be suffixed with a type, (optionally provideasciias a choice type likestd::ascii::Char). Changestrtocharsandasciitoasciis. Preferably rust would support this directly, otherwise do transmutions or similar at some point. Also introduceascii-to-char,asciis-to-charsand the inverse operations which returnopt -
add ascii operations like
fn Char-if-ascii-to-lower char : char fn Char-if-ascii-to-upper char : char fn Char-is-ascii-lower char : Opt ascii fn Char-is-ascii-upper char : Opt ascii fn Ascii-rid fn Ascii-dup fn Ascii-to-u32 fn Ascii-order fn Ascii-to-lower ascii : ascii fn Ascii-to-upper ascii : ascii fn Ascii-is-lower ascii : Opt ascii # maybe 'yes.'no. instead fn Ascii-is-upper ascii : Opt ascii # maybe 'yes.'no. instead -
add
Buf-opt-span-add-repeat,Buf-span-add-repeat,Buf-opt-span-add-repeat-length-positive, maybe even unfold -
add
Range-step .start u32 .length p32,Opt-range-step,(Opt-)Range-dup,(Opt-)Range-rid, probably also(Opt-)Range-elongate,(Opt-)Span-range -
combine scc stuff into the parser state to avoid walking the whole AST for info we could already have collected. Comes at the cost of a thicker ParseState, probably still worth. For extra convenience, it may be reasonable to implement some ByteDecode and ByteEncode traits in rust directly, so that in the common case that the state type is fully known you can hot reload with close to no glue code
-
add byte-level APIs, like
Buf-opt-span-take-i32 eniannessandBuf-opt-span-take-f32 enianness. Ultimately, these sould allow got reloading or simple byte protocols in general -
(probably not that good of an idea) to the above effect, it could be nicer to add ultra-basic macro support, so e.g.
!u32 "3"whereu32is of type_fn str, 'success u32 'failure str(instead of3 u32) which would evaluate the given function (which should return'error str (?) 'ok Value). This would allow userland to create e.g. hex parsing functions, arabic number systems, string raw bytes stuff etc. -
(probably not that good of an idea) consider not counting function calls as using up a function variable. The disadvantage is that "overplacing" a variable step-wise doesn't work anymore if not wrapped somehow. Maybe a small price to pay And where no
Lengthfield can be instantiated -
consider replacing kebab-case with camelCase/PascalCase. while I do much prefer the typing experience of kebab-case, camelCase is shorter (!!), think
BufOptSpanAddStrCharscompared toBuf-opt-span-add-str-chars(5 chars less, 20%!) and potentially more readable (?) due to clearer distinction to _ and (this won't matter as much if call and construct syntax does not involve _). Take a bigger example, convert the case and see how it feels -
(not fully sure) Add explicit field punning syntax: Add pattern syntax
_(untyped) /_ value-type(typed) (and maybe expression syntax_) where_behaves like a variable with the name of the parent. So e.g..field (_ value-type)would introduce a variable namedfield. and pattern'variant _would introduce a variable namedvariant. Likewise,linked-list-cons .nodes _ .linked-list numbers .new 3 u32would work if a variable namednodesexists. If no parent name exists, an error is thrown. The only goal here is making record patterns more convenient to work with (similar to swifts named parameters). The biggest worry I have is name clashes. Things like rename might also become a little more complex. I would normally not consider this as a feature, but since sloe is so painfully explicit, I feel users deserve some sugar for their effort. -
add field spread syntax for types where overlapping field names is okay as long as their value types are equal
-
add variant spread syntax
''existing-choice-type 'other-variants-before-and-or-after ...(only in types) analogue to the field spread syntax -
when checking, avoid shortcutting early when possible, still traversing sub-items even when a clear error has been found
-
add c# or swift or erlang compilation as well if there is demand
-
add source maps for mjs
-
verify that origin creation is correct for all kinds of recursion! e.g. this one seems on the edge of correct: different bufs have the same origin but their slots can't intermix.
fn Recurse .consume-origin consume-origin Origin _consume-origin, . .result-origin result-origin _result-origin : Buf _result-origin, u32 = ^ local-origin ? Buf-empty{u32} consume-origin [temporary] ? Recurse local-origin result-origin [result] ? Buf-add .buf temporary .new 1 u32 [.slot slot .buf temporary] ... resultIf we find a problem, creating a new
originshould be disallowed in (mutually) recursive calls. This is a bit restrictive but alright I believe. If feeling motived, look into proof languages and make sure this is rock solid -
consider a more general API for origins to be useful outside of giving them to new Bufs. To do that, allow origins to issue
Origin-uses. Questions:- is there a use for this or is it only for type gymnasts?
- is this even properly encapsulatable = useful considering that sloe only allows constructing structural types?
Like, if Buf was represented as
.array ... .origin Origin ...and Slot as.index u32 .
-
improve memory efficiency of string operations (currently buf of char). This is probably inefficient because:
- more work on program boundaries. E.g. instead of validating data, then reusing the bytes, we need to re-allocate them and then finally un-convert them into utf-8 anyway
- most bytes are 3/4th 0s because ascii is so common. wasted space is bad for the cache and memory usage
If these somehow turn out to be nonconcerns (e.g. through array-of-union(enum) optimizations) that would be cool as well since
Buf _, charis a much nicer API to work with
-
(not possible with origin-erased probably) zig-only: store an allocator within an origin (but! what about unset_slice? That one should probably store an allocator, too, and re-allocate if the buf origin allocator reference differs. I think this can be slightly unintuitive for sloe users but should in practice be okay). This achieves that origins created from within sloe code are arena-allocated and origins from user code are (usually) not, choosing e.g. MemoryPool.Aligned (does that actually work even?)
-
switch to a symmetric unerase API which enforces that all values inside are unerased. This would allow
Buf-origin-erasedto be removed in favor ofBuf (Origin erased, _part)andOrigin-erased-rid/Origin-erased-mapto just work. To that end, introducefn Origin-unerase .erased Origin-erased .origin Origin _o : Origin-isolated _o fn Origin-isolated-split Origin-isolated _o, .a _a .b _b : .a Origin-isolated _o, _a .b origin-isolated _o, _bbut I couldn't find something reasonable for choice types
-
(once there is an easy way to check if a pointer is aligned in rust) change
cast_or_rid_and_allocateto recover alignment differences if the address happens to align -
(once allocator API is stabilized) allocate all collections with an origin that was declared in sloe using a locally-passed
impl Allocator<> -
I think in theory there should be all the bits and pieces present to allow for struct-of-arrays and arrays-of-variant-values (made up name). E.g. internally compiling
Buf _origin, .a A .b Bto.a Buf _a-origin, A .b Buf _b-origin, BBuf _origin, 'a A 'b Btoor.slots Buf _slots-origin, 'a Slot a-origin 'b Slot b-origin .a Buf _a-origin, A .b Buf _b-origin, B.tags Buf _tags-origin, 'a . 'b . .slots Buf _slots-origin, Slot ??-origin .a Buf _a-origin, A .b Buf _b-origin, B
-
look into
soa_derivefor rust, maybe this already does most of the useful work -
(very out of scope but thinking never hurts) imagine what a logic programming language with this concept would look like. I imagine it wouldn't look much different (!) though with some different tradeoffs (e.g. more complex stdlib and compiler output, potentially a different typing and exhaustivess system)
As a hobby language that deliberately cannot by itself interface with the operating system, C etc. we can afford to skip many complex features. First some smaller-scale rejected ideas
-
temporary, unrestricted sharing. One might imagine that this could be implemented in API land with something like
ty Immutable _origin, _value ty Shared _origin, _value fn To-immutable .origin Origin _origin, _part .value _value : Immutable (Origin _origin, _part) _value fn Immutable-share Immutable _origin, _value : .immutable Immutable _origin, _value .shared Shared _origin, _value fn Shared-dup Shared _origin, _value fn Shared-rid Shared _origin, _value : . fn To-mutable Immutable _origin, _value : _valuecombined with various APIs to e.g.
Shared-span-lengthwhich take and give backImmutable. However, this is horrible!- Shared is entirely useless (since it isn't possible to e.g. define map, merge etc.). The only vaguely plausible utility of
Immutableis preventing mutation - it doesn't mix at all with non-shared functions and types
- wrapping and unwrapping Shared types is a giant pain
- Shared is entirely useless (since it isn't possible to e.g. define map, merge etc.). The only vaguely plausible utility of
-
(rejected because not useful enough) add
Outside _value,Origin-isolate(which wraps into anOutside),Outside-origin-unisolate(which unwraps anOutside),Outside(which wraps into anOutside). This API would allow any value to be transported through the lifecycle of anOrigin-erased(e.g. a large enum or complex copy-able record) without any hassle.Strictly speaking, this is less useful than just unisolating it value by value but it's also much more convenient.
Optionally, it could be possible to construct, map, pair, unpair such a value in sloe (also: read primitive values from it). It just doesn't seem any useful.
-
rename
Origin-isolated _origin, _valuetoIn _origin, _valuefor brevity. Then change (while keeping all existing operations) for consistencyOrigin _origin, _partto ``Slot Origin _origin, _parttoIn _origin, Slot _partSpan Origin _origin, _parttoIn _origin, Span _partBuf (Origin _origin, _part), _itemtoIn _origin, Buf _part, _item
However, this currently does not work because e.g. most Array operations parameterize the whole origin (which includes the part). If sloe didn't have an origin part system, this could be much simpler. The part system is only used for Origin-isolate (where e.g. A buf could store Slots to different origins). If it was possible to get the same functionality with other API means, I'd be all for it. However, I think this is straight impossible because it would be impossible* to find out e.g. what Origin a Slot was initially referencing into :/
- I think it is vaguely possible by assigning part names like
.a .and.b .when merging multipleIns with different origins. But this all seems very very hairy
-
(soft reject) allow expressions whose type is known (basically anything except inputs to queries) to omit extra type info (namely number, 'variant{} and project-fn{}). I'm a little torn because this makes construction inconsistent and increases the distance between the known type and expression. On the other hand this is already the case for query case patterns (deliberately so) but has a much higher convenience gain there. All in all, I think e.g. only allowing types in the first query case result / first array item etc. is more trouble than is worth. But it's not a clear-cut call
-
add special syntax
fn-oncethat automatically assembles the environment from the used local variables. Rejected in favor of more explicit construction with contextual names and potentially multiple fns. More info in "not coherently formulated thoughts" -
add tuples: (* a * b * c). I dislike them conceptually but operations like
U32-addorU32-dupare nicer with them. The field names.a .bare just noise. These can also be used for theArraytype paramerer. Adding tuples might make for a nicer user interface when calling from rust -
consider allowing
origin nameat the project scope. This allows reducing the number of type parameters flying around in things likeExpression _expressions, _patterns, _types, _source, _cases, ...if desired. It also makes initial_state much easier to call from the rust side (though we need to be careful how...). Rejected because this makes it more or less impossible to run multiple sloe instances from a single rust program Issue is that in general single-return-continuation is rare in sloe -
requiring all (!) generic type parameters to be passed to calls I feel like this is more "natural", easier to type-check but way more verbose / redundant. And especially because having many origin type variables is common, this sadly won't fly
-
adding function call syntax sugar similar to piping. While this is bloody wonderful (succinct, intuitive-ish, great for builders), it doesn't quite have much of a purpose which pattern matching doesn't fill well already. But more importantly it is quite limiting (requires positional arguments, requires them in the right order, doesn't apply to variants and similar). It also introduces "yet another way of writing the same code" which is dislike
-
(rejected, but interesting in theory) making
Bufetc store multiple kinds of data (heterogenous) and letting them give outSlot origin, data-typeandSpan origin, data-type. This means that usually only oneoriginneeds to be passed to things likeexpressionand slots/spans actually tell you what data they point to. Similarly, only one buf needs to be passed around. This makes the porpose ofBufbeing allocator-ish spaces rather that collections to query and edit more clear and makes passing them around to operations very simple, e.g.expression-end .expression Expression _origin .data Buf _origin, ... : .buf Buf _origin ... .end text-position. This would also enable below representation of tagged unions (this representation is not very flexible and otherwise also contradicts other basics of sloe):ty Expression-slot _origin 'int Slot _origin, i32 'plus Slot _origin, .left Expression-slot _origin .right Expression-slot _origin ...This also means slices etc need to be stored separately in the origin buf. The issue currently is that it feels hard to optimally construct/query such a heterogenous structure. Its structure must be created at compile-time. Dynamically this doesn't fly:
Buf origin = { bucket: Map<for type_byte_size: { key: type_byte_size, value: Buf<type_byte_size> }> }. However, really providing this in sloe would require sloe to add some kind of "type variable must be record" constraint:^ buf-origin ? Buf-empty{.expression expression .pattern Pattern buf-origin} buf-origin [buf] ? Buf-add .buf buf .new some-expression [.buf buf slot some-expression-slot] ? Buf-add .buf buf .new some-pattern [.buf buf .slot some-pattern-slot] ...This is probably doable in zig but hardly in rust without significant macro magic. Any ideas welcome!
-
allowing
.. ('variant ...)with a single variant and untyped variant expressions. No, should consistently use single-field record -
field and variants are changed so field names and variant names are uppercase and
.is spread (same for'), e.g.ty event 'Counter-clicked . 'Mouse-moved .X u32 .Y u32The benefit is that the question above is answered (single field = single variant). Overall this is "more correct" than the current solution. Rejected because this is harder to type (and would require a change of type variable syntax)
-
(rejection not final for all eternity. If you have a good use case, I'll support it) allow field and variant names to start with digit, upper-case and -, like
fn Char-dup char char : .0 char .1 char. One nice thing is that this matches what most language use as field names for tuples. This is also a little bit confusing but you don't have to use it. Use cases are e.g.ty bit '0 . '1 .,type board-pin '0 . '1 . '3 . '10 .and nicer array records. Not included currently for consistency and simplicity. -
switch from error{OutOfMemory}! to anyerror! for ease of use with external functions. Rejected because zig errors should be explicitly handled by sloe
-
rename
CalltoRun. "call" only makes sense if you've already heard it in that context, no?Hmm. Actually, considering literally all of programming uses "call", including for technical terms (call site, calling convention, tail-call elimination, etc.) users are way, way more likely to expect the name
Call. Providing both names is not an option for consistency. Sloe is also not targetting absolute programming beginners, so... should be fine as is.
While seemingly convenient and magnitudes better than regular mutable pointers,
- it's less obvious than passing values through
- there's sometimes no easy way to change the name of a resulting value that represents something different now
- there's no way to "reconstruct" a different out value. Especially for non-trivial edits the &mut approach can get messy or it's straight up impossible and parts will need to get cloned unnecessarily
- there's no way to change the type (e.g. from
Opt SpantoSpan) - there's two ways to specify most conversions, with usually no clear method of converting one to the other
- it's surprisingly common that one path consumes an argument, the other path keeps it in tact (e.g. when searching a tree with intermediate information. Either we find something, consuming the context or we come up empty-handed with the original context, like
fn .context context ... : 'done found 'going contextwhere found contains some parts of the context). This isn't modelled well with&mut &mutmeans the resulting changed collection is not returned, making use as the input to another function impossible. This almost necessarily results in the classic procedural-style statement form as opposed to the functional-style expression form. Minor gripe: especially in languages that don't allow local scopes with local returns (far, far too many) this basically makes it impossible to locally introduce a value, change it and implant it somewhere; instead you have to move the variable up to the top level.- returning
.(like returningUnitin gleam) feels super awkward to my brain. Most often, languages then automatically return void/... in the absence of a return and introduce all kinds of constructs like re-assignable variables, additional constructs for looping and branching that all can only return void/... . To my brain, this just confuses matters; it loves simple to follow flow of state! - &mut usually comes with the need to check for non-overlapping references to the same parts of data. This isn't possible with owned data passing in the first place
- &mut usually necessitates the need for offering the same APIs in two shapes, e.g.
make_uppercase(&mut self)vsto_uppercase(self)->Selfon rust's char/str types,take()/take_mut(),std::mem::swapetc. This to me just feels wrong - it's hard to be precise in what parts of a type are captured exclusively, shared or unused
rusts immutable references & have some similar trade-offs but seem kind of unavoidable at least for languages like rust.
- its type cannot be specified. as such, it cannot generally be stored as part of a type
- no clear unified interface. There could be multiple functions, there could be an output that additionally returns the captured values (allowing it to be called again), there could be an output that only sometimes returns the captured variables, etc.
- I personally never had a need for this. Usually you can just make the environment a type variable and you're golden
I'm strangely really convinced that this is the obvious, correct design decision (for most programming languages at that!).
Note that the current design does not natively have a dyn Fn; it needs to be manually emulated via an explicit choice type.
- traits introduce a crazy amount of complexity
- If really necessary, traits can be represented using arguments. I have yet to hit any complexities with this.
- attaching a set of functions to one "subject" seems super strange to me. Operations usually take different objects and create something new
- traits create a "one-fits-all" interface. Thinking to e.g. rusts
clone(&),drop(&mut)orto_string(&), they seem super sensible in concept but fall short when trying to clone into a different allocator, trying to use a different string representation, or trying to mutate some backing storage on drop. These restrictions can sometimes kind of be circumvented in annoying ways. It all just seems so arbitrary for no apparent reason. - traits are usually touted as a sensible solution for operator overloading. sloe does not have operators
- traits push languages in the direction of nominal record and choice types. This isn't wrong per se but to me they usually start to feel clunky to use and I then use them less often then would be helpful (e.g. for long parameter lists or multiple outputs)
Because traits cover a vast theoretical area of use, they tend to be used a bunch. I've never found them particularily pleasant to use. Libraries often only expose some functionality through these, without proper documentation. Incidentally, I've also found editor tooling to be lacking in these areas, not knowing if you want to look at the general or specific function.
- operators introduce a good amount of complexity: infix (and prefix) notation, associativity, precedence, most likely a way to overload based on context
- edge-case behavior (e.g. saturating vs overflowing vs checked vs carry vs ...) should be easier to control
- in general, operators are concise but as a result quite ambiguous. For example, changing a boolean to an integer may silently not generate a compiler error when
!is binary not, or when changing an int to a string with+ - while numbers, bool and bit operations are not that uncommon, there are features that would deserve these symbols more, even in typical imparative languages (think
return,switch { case },structure,import,public,static,void,null, ...) - allowing infix
-and prefix-leads can lead to very confusing situations likecall-1but more importantly using-as an operator pretty much prevents languages from using the superior (easier to type) kebab-style for identifiers
Somehow despite it's issues (math syntax kind of sucks, even the tiny subset), operators are one of the most prevalent features in programming languages, even hobby and experimental ones (0th class citizen).
Sloe had positional arguments once. It's the more pracical and convenient choice, and makes interfacing with rust/zig/js simpler.
The decision to remove them is largely personal. I get lost easily in long argument lists and I have have no others users to please ^^. My (bad) more objective arguments are:
- it's tough (usually) to annotate a function whose arguments and argument types are unknown. E.g. what would
Fn-dup's type be? - positional arguments (usually) means no passing in bulk
fn U32-square-clamp natural u32 : u32 = U32-add-clamp U32-dup natural
Today "positionality" in general is pretty much absent in sloe (except for type parameters). E.g. positional arguments are super convenient, so they tend to be used for everything, even arguments that would benefit from a clear description.
Features I've added which are formally fully replacible by other existing features. If you're looking to learn from sloe's central ideas, maybe do not learn from these:
-
record spread. It provides an alternative syntax sugar for something that could already be expressed. I originally introduced it to make builders like string builders less jarring but I'm not fully convinced this direction worked (e.g. maybe adding extra syntax for field punning would have been more explicit and just as concise?). Especially for query case patterns where only one spread can exist per pattern, I took a very long time before changing my mind to add it. It enables the "use the defaults except" pattern which would be inpossibly annoying otherwise:
Some-fn ? Some-fn-defaults [.. all .except except] ? Except-rid except [.] .. all .except new-valueI've changed my mind on this being okay because you need to handle all fields anyway. It's one of those "only need it in 5% of cases but then its unreplaceable" features - the nightmare of a language designer
-
nested pattern matching. It's existence makes compilation, exhaustiveness-checking, error messages and the possibility of flow-typing-like matching (e.g. matching 'a in 'a'b'c leaving 'b'c) a bit harder. It also creates a "two modes of matching" problem: You e.g. can't match on numbers, chars, strings, span start and lengths etc. And so you sometimes need an extra step, leading to nested matches anyway (does not feel consistent). It also "takes control from the user into the magic hands of the compiler" and thus it may run checks etc. in a different order than you have. I originally introduced it to make e.g. matching on multiple
Opts easier. It helps keep context clear and visible like "if the left sub is empty and the right sub is a branch with an empty left side, do this". Honestly I should not have been so hasty to add this feature -
stack-allocated array syntax. It provides an alternative syntax sugar for something that could already be expressed as repeated queried function calls. I'm convinced that a feature like this would be very asked for if it didn't exist. Adding a bulk of items to a buf seems very useful on first sight because
- all kinds of examples and tests start with manually adding items. Doing this one by one seems like cringe busywork.
- building uis or any kind of trees programmatically, you more than often end up needing to specify sub-nodes of a parent. Adding this bulk of nodes as an array is only natural and avoids so much noise.
Well, what are the alternatives, then?
- simply provide
Buf-opt-span-add2/3/4/5/6/7/8etc. While it really doesn't feel good, it's not very far from solutions of production languages, see e.g. java'slist.of2/3/4/.... (Though adding or removing items means adjusting the number which is annoying, especially for the argument field names) - provide and suggest better primitives and helpers. For example, instead of providing a list of modifiers, it may just make sense to e.g. provide a record of options or use builder-style helpers for the individual properties
- introduce syntactical diabetis or macro-esque bullshit for repeated function calls
(the above is obviously insane in a bad way, but there may be a middle-ground)
... ? Stack-cons* .. example-stack .new 39 u32 [-]? * ..- .new 3 u32 [-]? * ..- .new 6 u32 [-]? * ..- .new 9 u32
So yeah, these aren't amazing either.
I'd say domains where languages like performance-aware safe rust/C#/swift/go stand today:
- not extensive enough to have a front seat in systems programming, but comfortably sitting on top of a somewhat thin platform layer
- not as easy to use as scripting languages like python, gleam, lua, elm, prolog, etc.
- mainly used for the subset of tools, applications or similar where maintainability, robustness and being easy to reason about is more important than dev speed
That's far from general-purpose! Don't be afraid to program in a language sloe compiles to for tasks sloe feels annoying to use for. E.g. I imagine writing a recursive file watcher in sloe is not fun, so just "outsource" it :)
- you can use a style which prominently uses
Buf-pre-allocate-at-leastwhich cleanly fails. I think this is a reasonable compromise because pre-allocating is a very useful and prevalent pattern anyway when running out of memory is possible. - sloe is already too tedious. I certainly would hate (if it was the default)
- sloe's out of memory handling is already relatively graceful. E.g. when outputting zig, functions will return an explicit error.OutOfMemory. In rust, panicing on failed allocation is safe and the default.
- output language targets like js do not support this anyway
The best user experience interfacing with sloe code from existing (system-level) languages is directly generating code in that language. Just sharing type names, structs, tagged unions, function signatures etc without any work by you is tasty enough. And if/once you outgrow sloe, you have all the code right there (that's the hope anyway but output readability is likely way wose than as if it was hand-written). Being easy to transpile is an explicit goal of sloe, enabled by its very limited set of features.
It did that before and it does it's job. I imagine the current style leaves some performance on the table but I'd be surprised if it was too slow for its only potential temporary user, the human reading this (<3).
-
consider renaming
Origin-isolated-mergeto*-pairand*-splitto*-unpair -
add
Opaque _value(maybe should be "foreign" instead),Opaque-origin-isolate,Opaque-origin-unisolate.Opaque _valueis compiled to_valuebut is inaccessible in sloe.Important! opaque_origin_isolate/_unisolate should not be public in the output code as it's unsafe to use from the output language
-
another issue in this space: creating a partial Origin-erased with the erased value using
Origin erased, .but from an origin that has a part, likeOrigin some-origin, .some-part .. -
consider adding
Buf-step,Buf-map-or-rid-and-allocate,Buf-set-count, maybeBuf-combined-set-unset-length. The first 2 enable "spooky action at a distance" andBuf-(opt-)span-*operations should still be prefered if possible. However, adding them is necessary to enable more "data-oriented design" (less jumping around) and to make buf handling less painful. The only real reservation I have about this is that is is mutually exclusive to anUnset-slot/Unset-spanAPI (which I have deliberately removed but it still hurts to have let it go).
In rust, collections tend to own their item data, so safely keeping references reaching inside is tough.
Alternatively, we could reach for Range<usize> and usize but we've lost ties to the origin structure and rust does not (yet?) have a mechanism for temporarily assuming actual ownership over some part of a parent structure.
This relationship is flipped on it's head in sloe: All items of collections are divided into slots and spans which are owned by the code that parked values there in the first place.
Honestly this idea seems "obviously" useful and it's surprising I can't find other languages that lean into it (there is rust which at least enables it in userland). I assume one reason is that linear types are required in some part to avoid leaks all over the place.
One way this helps is that nested collections aren't segmented: what is usually Buf<Box<str>> aka n separate memory pieces can be e.g. Buf ... Span str-origin + Str str-origin
(in rust there are I think crates like oroborus for this)
since each variable can be used at most once, most introduced names that would traditionally be considered "shadowed" are aready out of scope in sloe. When their scopes actually overlap though, you'll get an error
I love how linear types somewhat mirror the functionality of defer ...getRidOfIt(); but without the yucky control flow. All operations happen in the specified order in sloe!
This also simplifies code generation
When I started imagining this language I believed the few core concepts to be pretty unique. Reading more on the various aspects, it turns out this cake was already in the oven twice. For example, using indexes that are marked to uniquely reference their origin array at compile time seems to have been individually already explored by many cool people. Many modern languages seem to converge to a similar (or even better) design (carbon, visions for rust, valen, dada, various libraries, zig).
Now I would indeed say that maybe there was never a place or future for sloe but exploring these hot topics and arriving at a similar place as many others was still nice to learn :---)