Aliasing-xor-mutation WITHOUT borrow-checker

Thank you for the links. Unfortunately info there is extremely shallow. I won't elaborate here though as this isn't Mojo forum.

You are right about some edge cases. I just haven't encountered cases where I need to annotate a lifetime yet. For example

def abc(input: ref String) -> ref String:
    return input

That works without a lifetime annotation. It works in Rust too, but I haven't encountered code where I need to annotate a lifetime in Mojo so far, maybe I would when I code more complex stuff

I wrote it incorrectly. I wanted to mention that Mojo doesn't need to wait until the end of the scope to free memory, it returns memory faster to the allocator with no stall time. It does that by doing data flow analysis, if no other subsequent code uses the variable, the Mojo compiler inserts a memory freeing call. This is very good because take a look at this code:

// scope
{
    // many lines of code

    // 1GB memory
    let buffer = processing;
    save_to_db(buffer);

    // this processing step would have less free RAM available
    // unless the user drops the buffer manually, which I very rarely see Rust users do
    // because Rust would not drop it despite nobody else using it anymore
    // since the end of the scope hasn't been reached yet
    // so that memory becomes stale
    long_processing(...)

    // many lines of code

// end of scope
}

Whereas in Mojo, all memory that is no longer used will be immediately returned to the allocator, so subsequent code has more free memory available, since it doesn't need to wait for the end of the scope. This technique would make Rust programs smoother because no stale memory if it were added to Rust

Does Mojo allow running arbitrary code in destructors? Doing this in Rust would cause the runtime behavior of programs (not limited to resource consumption, but actual correctness like for locks) to depend on at exactly which point the borrow checker believes it safe to destroy a variable. This would both be surprising and would mean that improving the borrow checker would be a breaking change. We did change drop order a while back to fix some edge cases, but we had to do this over an edition boundary. We use the same borrow checker across all editions (NLL was initially only introduced for crates that used the 2018 edition with the 2015 edition using migration mode, but that was just a temporary thing to make the transition smoother: Non-lexical lifetimes (NLL) fully stable | Rust Blog) and there is no way to ensure the expansion of macros runs with the right borrow checker for an edition as you can't mix borrow checkers within a single function.

Yes, you can write a custom del filled with custom code, just like writing a custom drop in Rust. In Mojo, specifically for parts that run better with RAII, it is still kept as RAII, using the with keyword like this (it's called a context manager):

var mutex = Mutex(0)

with mutex.lock() as data:
    data += 1
    # mutex unlocks at end of the context manager

Documentation: Errors, error handling, and context managers | Mojo

The eager drop technique can be implemented without breaking existing code by applying it in a fine grained manner. For example:

// specific to this struct only
#[eager_drop]
struct ABC
impl Drop for ABC { ... }

// specific to this crate only
// main.rs or lib.rs
#![use(eager_drop)]

Perhaps you could bring some examples where Mojo doesn't need lifetimes but Rust does? Right now your claim is supported by neither examples nor by Mojo's documentation.

Besides bjorn3's point, this has nothing to do with the borrow checker or aliasing xor mutability.

2 Likes

This is code that requires a lifetime annotation in Rust, but does not require one in Mojo:

fn main() {
    let a = abc(&String::from("abc"), &String::from("abcd"));
    println!("{}", a);
}

fn abc<'a>(a: &'a String, b: &'a String) -> &'a String {
    if a.len() < b.len() {
        a
    } else {
        b
    }
}

Mojo:

def main():
    var a = abc("abc", "abcd")
    print(a)

def abc(a: String, b: String) -> String:
    if a.byte_length() < b.byte_length():
        return a
    else:
        return b

It has something to do with the borrow checker. For example this example code will not compile in Rust because of its scope based drop:

fn main() {
    abc(1, 2); 
}

fn abc(a: i32, b: i32) {
    let mut buffer = Vec::new();
    if a < b {
        let a = String::from("abc");
        buffer.push(&a);
    } else {
        let a = String::from("abcd");
        buffer.push(&a);
    }
    
    println!("{:?}", buffer[0]);
}

But it compiles in Mojo because of its data flow analysis:

def main():
    abc(1, 2)
    
def abc(a: Int, b: Int):
    var buffer = List[String]()
    if a < b:
        var a = "abc"
        ref b = a
        buffer.append(b)
    else:
        var a = "abcd"       
        ref b = a
        buffer.append(b)

    print(buffer[0])

I personally like that the Rust version requires explicitly-annotated lifetimes here, so that 1) if I'm maintaining the code, the compiler can give me an error if I change the body in a way that breaks the existing lifetimes, or 2) if I'm using the code, I can tell the lifetimes by looking at the signature, instead of having to look into the body to figure out what the lifetimes are.

2 Likes

I don't like that, and I read lifetime is one of the biggest problem that user does not want to deal with. People in Rust mostly just say to clone it, as a way to avoid lifetime annotations, which sacrifices performance. Mojo can also give compile errors because you can not violate its lifetimes. You are misunderstanding it, the difference is specifically in how parameters work, in Rust the default is move, in Mojo it is the opposite, the default is pass by reference, no need to look at anything else too, if it is a plain type, it is an immutable reference. If there is a var keyword, it is a value. If there is a mut keyword, it is a mutable reference. Except for primitive types like int and others which are copied because it is faster. That default pass by reference combined with data flow analysis is what cures a lot of the lifetime pain

In short, memory cleanup in Mojo works naturally, the same way manual memory cleanup works. In C, if memory is never used again in subsequent code, the free function is called. Well, Mojo does exactly that

So I see the following mojo code:

def abc(a: String, b: String) -> String:
    # Very long body I don't want to read through

How do I know how long the returned String reference lasts?

In Rust, I can easily tell from the signature:

// Lasts as long as both input references.
fn abc<'a>(a: &'a String, b: &'a String) -> &'a String;
// Lasts as long as the first input reference, regardless of the second.
fn abc<'a>(a: &'a String, b: &String) -> &'a String;
// Lasts as long as the second input reference, regardless of the first.
fn abc<'a>(a: &String, b: &'a String) -> &'a String;
// Lasts as long as it is kept around, regardless of either input.
fn abc(a: &String, b: &String) -> &'static String;

How does Mojo distinguish between those different cases without lifetimes in the method signatures? If the answer is "it does data flow analysis on the body of the function to infer the lifetimes", then I refer you to my previous comment for why that's undesirable.

3 Likes

That's the caller's business, not the function's. In manual memory management (C, C++), the function returns a value/pointer and the caller decides how long to keep it alive. The function doesn't dictate that. Rust's lifetime annotations are the function imposing its internal structure onto the caller's code

"How do I know how long it lasts?"

You don't need to, and you shouldn't have to. The value lives as long as you use it. That's how every other language works. The compiler tracks data flow, if you stop using it, it gets cleaned up. If you keep using it, it stays alive. To understand it, it is like everytime Rust user fix borrow error using Rc or Arc, but zero cost. The function has no business dictating that

Your four examples actually demonstrate the problem, in example 1, you're forcing the caller to keep both a and b alive even if the body only touches a. That's implementation detail leaking into the public API. If you later refactor the body to only use b, the signature must change, and every caller breaks, even though the code is still perfectly safe. A return value should live as long as the caller needs it, not as long as the function's internals happen to require. That's not a feature, that's backwards encapsulation

Can you explain this code's behavior:

def manifest() -> String:
    var a = "abc"
    return abc(a)

def abc(a: String) -> String:
    return "def"

Rust can do this because abc and manifest would return &'static str. But in Mojo, if abc starts returning a instead, how is manifest going to ensure that its return value lives long enough? The callers now need to call a destructor as well. Is a moved to the heap? If the callers somehow know that it is not 'static, where is "the destructor needs called" stored?

1 Like

To understand the Mojo model, you don't need to learn anything special, because it works the same way as every other programming language (whether GC or manual memory management), once memory is no longer used it is freed, and if it is still being used, for example passed into another variable or into a function, it is freed after that last user finishes, which is naturally the correct memory behavior of freeing it after it is no longer needed, not freeing after the end of scope. Since var a in the abc function is no longer used and abc returns a different value, var a is freed inside the abc function. In Mojo it is very easy, the code adapts to you, not you adapting to rigid rules. So when you return var a from the abc() function, the value is now being used, so it stays alive for as long as the caller uses the return value, as soon as the caller stops using it, it will be freed. The callee has no business ensuring and dictating how long the return value lives. It is the caller that decides. If you use the return value:

var a = abc()

ref b = a

fn(a)

fn2(b)

As long as the value is still being used, it will automatically stay alive. As soon as the value stops being used:

// for example here I stop using everything that uses a, so a is automatically freed right after the last usage which is fn2()

The 'static is itself a lifetime annotation leaking implementation detail into the signature. The caller now must know the value is static. If you later change the implementation to return a dynamically constructed string, the signature breaks and every caller breaks with it. That's the encapsulation violation

Am I crazy or is this an apples to oranges comparison? It looks to me like the Mojo String type is an owning type much like Rust's String. So to make this an equivalent comparison you must either change Rust to use String which eliminates lifetimes altogether, or you have to change the Mojo code to use StringSpan or have refs to the strings (I'm not sure what exactly that would look like).

2 Likes

Here's the Rust equivalent of the Mojo code:

fn abc(a: &String, b: &String) -> String {
    if a.len() < b.len() {
        a.into()
    } else {
        b.into()
    }
}

Notice how it returns String and it's doing an additional .into() conversion from &String to String? That's a hidden allocation+copy in Mojo! Moreover notice how we also don't need lifetimes in Rust.

IMO it's clear here how the more implicit things you want the code to be, the easier it is to hit footguns like these.

Edit: in case it wasn't clear, in Mojo parameters default to references, but return types default to owned types, and the returned values can be implicitly converted to the return type (Functions | Mojo)

This is a function directly from Mojo docs, showing that you do need annotations when returning references. You may argue the annotation is nicer than Rust's, but you still need one.

def pick_one(cond: Bool, ref a: String, ref b: String) -> ref[a, b] String:
    return a if cond else b 

This has another implicit copy. Your List is holding owned Strings, not references, so you're converting them somewhere. It's just yet another footgun.

1 Like

You are right about this one. At first I thought the default pass by reference applied to everything, but the correct behavior is that pass by reference only applies to parameters. For return types, the default is pass by value (copy or move depending on whether the type implements Move or not, with Move being the priority if it implements Move). So the code would be like this, still cleaner than Rust if you ask me

def main():
    var a = abc("abc", "abcd")
    print(a)

def abc(a: String, b: String) -> ref[a, b] String:
    if a.byte_length() < b.byte_length():
        return a
    else:
        return b

You are right about this one. At first I thought the default pass by reference applied to everything, but the correct behavior is that pass by reference only applies to parameters. For return types, the default is pass by value (copy or move depending on whether the type implements Move or not, with Move being the priority if it implements Move). So the code would be like this, still cleaner than Rust if you ask me

But you are also wrong. It does not copy the string. String is a heap type, and heap types in Mojo implement Movable, and Move is prioritized over Copy just like in Rust. In Rust, when passing by value it is either move or copy, if the type implements Move, Move is chosen. So it is a Move, a transfer of ownership, not a copy. That is zero cost, not a performance footgun

But it is still cleaner than Rust when the code is put side by side

Rust

fn main() {
    let a = abc(&String::from("abc"), &String::from("abcd"));
    println!("{}", a);
}

fn abc<'a>(a: &'a String, b: &'a String) -> &'a String {
    if a.len() < b.len() {
        a
    } else {
        b
    }
}

Mojo

def main():
    var a = abc("abc", "abcd")
    print(a)

def abc(a: String, b: String) -> ref[a, b] String:
    if a.byte_length() < b.byte_length():
        return a
    else:
        return b

In Rust keep repeating the lifetime over and over, there it appears 4 times. In Mojo just writing it naturally without repeating, only 2 times

You cannot move out of immutable references.

Your LLM is confusing you here, most likely because there are few examples of Mojo code/tutorials in its training data.

I hope you see however how you went from "Mojo doesn't need lifetime annotations" to "Mojo doesn't need lifetime annotations in this case" to "Mojo needs lifetimes annotations in this case otherwise it will do an implicit copy, but it still reads better than the Rust code".

All considered Mojo has the exact same underlying model as Rust, it has nothing revolutionary there. The differences come mostly from the different defaults and the choices of allowing some questionable implicit behaviours.

1 Like

So you are accusing me of using an LLM? Are you somehow mad? Because I am not using an LLM, LLMs can not write up to date Mojo code without being taught, they still know only the old syntax, for example they use fn, borrowed, etc. I was reading the Mojo lifetime origin documentation, which states that Mojo defaults to pass by reference. While the return value documentation is on a different page

Because the goal is to show what advantages other languages have as a reflection and to try to bring those advantages to Rust. There are still advantages that Rust does not have, namely all the performance benefits of ASAP (As Soon As Possible) drop

So every string is always heap allocated and there's no static storage optimization at all? Because the caller of manifest needs to know wheter the data is even allocated at all to know whether to destruct it. Or is everything a fat pointer carrying its own destruction function pointer around too?

If there is no static storage area for use, that's an…interesting design decision. I'm not sure it's worth not having lifetimes for the kinds of programs I've worked on. Some of those programs have been 30%+ runtime doing allocation, strlen, and deallocation of strings prior to plumbing static storage support through APIs.

Mojo has StaticString. Mojo exposes 3 string types to the user: String (heap), StringSpan (slice), and StaticString (&'static str). StaticString is a StringSpan that is static, it is a fat pointer, just like str in Rust, whether static or stack, it is a fat pointer

More detail : string_span | Mojo