首页 > AI前沿 > Reducing undefined behavior in the C language

Reducing undefined behavior in the C language

Hacker News 2026-10-09 10:02 3 阅读 查看原文
Reducing undefined behavior in the C language Why bother with C in 2026? It is, he said, still a great language. C is portable, stable over the long term, offers fast compilation, and the resulting binary code is fast. "What you see is what you get"; it is easy to look at C code and have some idea of what the computer will actually do. There are a lot of tools for working with the language, and C gets out of the way when necessary. The approach that was taken was to define the semantics of the language in terms of an abstract machine. All operations are to be executed as if they had run on that abstract machine, which may not exactly match the actual hardware. The observable behavior of the program must be what the abstract machine would have done. The "observable" part matters: access to volatile variables, being defined as observable, must happen exactly according to the abstract machine; everything else just has to produce the same eventual result. The standard gives a lot of freedom to compiler implementers; only the observable behavior has to be preserved. There are many aspects of that behavior that are either undefined or implementation-defined. These are not observable behavior, and thus do not constrain what compiler implementers can do. There are, of course, other specifications that can constrain compiler developers where the C standard does not; these include ABI requirements, standards like POSIX, or the need for backward compatibility. Undefined behavior comes about when a program does something that is either not portable or not defined by the standard at all. In such cases, the C89 standard states that it "imposes no requirements" on the implementation. Undefined behavior exists for a number of reasons. It allows implementations to support extensions, manage interactions with hardware-based safety mechanisms, and perform aggressive optimization, all while allowing difficult-to-detect errors to be ignored. It explicitly gives the compiler the right to ignore whole classes of hard-to-detect errors. Nasal demons The problem, Uecker said, is that the standard allows a compiler to do anything in response to undefined behavior, up to the point of invoking nasal demons. If a program contains any undefined behavior at all, according to compiler writers, then it has no expected semantics. The C++23 standard goes further to explicitly state that the standard imposes no requirements for these programs. That has led to widespread disagreements between developers about what can be expected from the language. For example, if you zero an entire structure (perhaps with a call to memset()), then write to specific fields, what will happen if you read from any padding bytes in that structure? Might they contain security-relevant data? A 2015 survey showed that there was no consensus on what should happen in that case. Or consider this simple code: extern int x; int f(int a, int b) { x = b ? 42 : 43; return a/b; } If b is zero, then the return statement is a division by zero, which is undefined behavior. In this case, is the compiler entitled to omit the test entirely and just execute x = 42? After all, the b = 0 case has no expected semantics, and can thus be ignored. There are compilers that will do exactly that. In the undefined-behavior case, the store to x is not observable behavior. But now consider this case: extern void g(int x); int f(int a, int b) { g(b ? 42 : 43); return a/b; } This might seem to be the same situation, with the compiler being entitled to remove the test and just pass 42 to g(), and some compilers have treated that way — but that compiler behavior was a bug. Imagine a definition of g() that calls exit() if b is zero. In that case, the division will never happen and the program's behavior is not undefined. So eliding the test and simply passing 42 to g() is incorrect. One more interesting case: volatile int x; int foo(int a, int b, bool store_to_x) { if (! store_to_x) return a/b; x = b; return a/b; } The question here is: can the compiler hoist the final division operation above assignment to x? If there are no semantics associated with the b = 0 case, then there is no change in observable behavior. This, too, is something compilers have done, but the C23 standard added a "no time travel" stipulation to disallow it. In C++, instead, hoisting must be explicitly prevented by inserting a call to std::observable_checkpoint(). Time-travel bugs should eventually go away, but there are a lot of other situations where, even if the standard is clear, compiler writers often disagree. These include reading of uninitialized variables (which is almost always defined), and equality comparisons of pointers, which is always defined, but is also miscompiled by both Clang and GCC. Fighting undefined behavior To try to address all of these problems and more, the C committee runs three study groups focused specifically on the memory object model, memory safety, and undefined behavior. There are currently about 100 instances of undefined behavior in the C standard, but the in-progress C2y draft has removed 45 of them. The situation is indeed getting better. There is an increasingly rich set of tools aimed at finding issues: compiler warnings, static analyzers, sanitizers, LLM-based tools, formal verification, and more. The number of situations where a compiler will emit a warning where possible undefined behavior is detected is growing; recent examples include better warnings for integer overflows and potential use-after-free situations. Static analyzers are available as standalone tools, but are also increasingly being built into the compilers themselves; GCC can now warn about a number of potential buffer-overflow situations, for example. Sanitizers work by inserting run-time checks; they can catch a lot of undefined behavior and, in trapping mode, be used for hardening as well. Memory safety has never been one of C's strong points, but Uecker wanted to make the point that it can be improved. That problem breaks down into three sub-problems: type safety, spatial memory safety, and temporal memory safety. C, he said, has a strong type system, and the remaining problems are fixable. Tagless unions, for example, can create type confusion, but the compiler can enforce types with some additional annotations. New diagnostics can catch unsafe casts from void. Type checking across translation units is traditionally not a huge problem in C, since header files are used to ensure consistent types, but the situation could be improved with a link-time checker. Spatial memory safety — bounds checking — is a partially solved problem; the compilers can perform array-bounds checking in many situations now. In some cases, some code changes are needed to fully benefit from this checking. Use of the counted_by attribute can enable checking for flexible array members, for example. Temporal memory safety — avoiding use-after-free bugs and the like — is harder, Uecker said, and Rust definitely has an advantage there. Still, better temporal memory-safety enforcement is possible. Architectures like CHERI can help here is well. Fil-C can find a lot of temporal-safety bugs. Can all of these tools and language changes get us to full memory safety? Completely solving the problem will require either expensive run-time checking or formal verification, he said. In the near future, the most complete results will be had with the combination of a restricted language and formal verification tools. Overall, he concluded, C is still a living language and is still improving. The C23 standard removed a number of problematic features, including old-style (K&R) function definitions, support for sign-magnitude and one's-complement machines, and trigraphs. It added bit-precise integer types, checked integer operations, and more. C2y will go further, adding case ranges, named for loops, the _Countof() macro to determine array lengths, and a lot of "demon removal". It will not achieve full memory safety for C, but that is an eventual possibility, and will become more practical over time. He ended by encouraging interested people to participate in the working groups. The video and slides from this talk are available. [Thanks to the Linux Foundation, LWN's travel sponsor, for supporting my travel for this event.] CPU-dependent behavior Posted Sep 28, 2026 17:27 UTC (Mon) by ballombe (subscriber, #9523) [Link] (45 responses) CPU-dependent behavior Posted Sep 28, 2026 17:34 UTC (Mon) by daroc (editor, #160859) [Link] (24 responses) The answer for Rust is that it depends on your compiler settings. By default, in debug builds it results in a panic which prints a stack-trace and exits. In release builds, it is guaranteed to wrap, which means it will evaluate to 0. But if you want to be certain, there are helper methods that will return an error explicitly if overflow would occur, for your program to handle however it likes. CPU-dependent behavior Posted Sep 28, 2026 18:35 UTC (Mon) by ballombe (subscriber, #9523) [Link] (13 responses) long fun(long a, long b) { return a>>b; } int main(void) { printf("%ld\n",fun(1,64)); } print 1 CPU-dependent behavior Posted Sep 28, 2026 18:48 UTC (Mon) by daroc (editor, #160859) [Link] (10 responses) CPU-dependent behavior Posted Sep 28, 2026 18:59 UTC (Mon) by ballombe (subscriber, #9523) [Link] CPU-dependent behavior Posted Sep 28, 2026 19:24 UTC (Mon) by kreijack (guest, #43513) [Link] (7 responses) For the SAL/SAR/SHL/SHR intel instructions [*], the counter register of the shift is masked with & 63 (or & 31 depending by the register width). So shifting by 64, is effectively shifting by 0: 1>>64 -> 1 >> (64 & 63) -> 1>>0 -> 1 The funny thing, is that testing this code in godbold, I got a lot of different results (0, 1, 0x7f49a68055c0 (!!!!!) ) depending by the combination of compiler/optimization. So, yes this is an undefined behavior, and we should expect any (un)reasonable result. [*] https://www.felixcloutier.com/x86/sal:sar:shl:shr /... The count operand can be an immediate value or the CL register. The count is masked to 5 bits (or 6 bits with a 64-bit operand). The count range is limited to 0 to 31 (or 63 with a 64-bit operand). A special opcode encoding is provided for a count of 1..../ CPU-dependent behavior Posted Sep 28, 2026 19:35 UTC (Mon) by ballombe (subscriber, #9523) [Link] (6 responses) CPU-dependent behavior Posted Sep 28, 2026 19:43 UTC (Mon) by mb (subscriber, #50428) [Link] (5 responses) It panics in debug mode and does whatever the processor does in release mode. Debug: playground::a: subq $40, %rsp movq %rdi, 8(%rsp) movq %rsi, 16(%rsp) movq %rdi, 24(%rsp) movq %rsi, 32(%rsp) cmpq $64, %rsi jae .LBB7_2 movq 8(%rsp), %rax movq 16(%rsp), %rcx andq $63, %rcx shrq %cl, %rax addq $40, %rsp retq .LBB7_2: leaq .Lanon.77e4319a821b47bb0a9368ddeddb101a.2(%rip), %rdi callq *core::panicking::panic_const::panic_const_shr_overflow@GOTPCREL(%rip) Release: playground::a: movq %rsi, %rcx movq %rdi, %rax shrq %cl, %rax retq CPU-dependent behavior Posted Sep 29, 2026 0:33 UTC (Tue) by walters (subscriber, #7396) [Link] While you probably know this, it's important to emphasize for the wider audience that Rust has almost equally ergonomic "checked" variants of arithmetic functions, in this case https://doc.rust-lang.org/stable/std/primitive.u64.html#m... that apply regardless of "debug" vs "release" builds - and it's often a good idea to use them. In fact, for most use cases they should probably be thought of as the default. CPU-dependent behavior Posted Sep 29, 2026 8:36 UTC (Tue) by ralfj (subscriber, #172874) [Link] (3 responses) It panics in debug mode and does whatever the processor does in release mode. It panics in debug mode and does whatever the processor does in release mode. Not quite. In release mode is always shifts by "offset & (bit_size - 1)". (To be even more pedantic, this is tied to -Cdebug-assertions, which is usually only set for debug builds but can be turned on in release builds as well.) The behavior of safe integer operations in Rust is fully deterministic and portable (modulo endianess). We do pay a small performance cost for that on some targets but we consider that worth it. Operations like unchecked_shl are available if you are in a hot loop and want to avoid the overhead of masking with bit_size - 1. CPU-dependent behavior Posted Sep 29, 2026 9:12 UTC (Tue) by ojeda (subscriber, #143370) [Link] (2 responses) -Cdebug-assertions -Cdebug-assertions More specifically, -Coverflow-checks, which I prefer because it is not uncommon to want to enable the overflow checks while not adding all the debug assertions, e.g. in Linux we enable the overflow checks by default at the moment but not the debug assertions (though it is possible we may need to switch that, which is why I requested -Coverflow-checks=report). CPU-dependent behavior Posted Sep 29, 2026 12:09 UTC (Tue) by ballombe (subscriber, #9523) [Link] (1 responses) CPU-dependent behavior Posted Sep 29, 2026 18:58 UTC (Tue) by ralfj (subscriber, #172874) [Link] CPU-dependent behavior Posted Sep 29, 2026 21:02 UTC (Tue) by stevie-oh (subscriber, #130795) [Link] It depends on the CPU because different CPUs have different ways of implementing dynamic shift operations. In the 6502 it could only shift one bit at a time. You'd literally jump into the middle of a shift/rotate function to process the number of bits. If you had a number that was out of range, it would execute random code. That's one of the reasons it's actually undefined behavior to do a negative shift or shift by more bits than the register can have. The there's the original 8080/8086/8088 chip, for example, which looked at the lower 8 bits and literally ran a loop. If you did "foo >> bar" with bar set to 255, it would loop for... (calculates) about 214 microseconds, mostly shifting zeroes -- but if you had bar set to 256, it wouldn't shift at all. And there's bad news: during those 214 microseconds, the CPU is completely locked up. The computer couldn't do anything else -- such as respond to time-sensitive hardware interrupts. On a multi-user server, that was bad. So the 286 put a limiter on it: it only honored the low 5 bits (since it only had 32-bit registers) which put a hard cap of 31 loops iterations. In newer machines, the microcode+circuitry that performs the shift(actually a rotation+mask) only processes the low 5-6 bits of the shift count field. After all, why would a CPU designer add the circuitry needed to handle a 7th bit when, under normal operation, it never gets used? This page shows a typical shift circuit *and* how the 386 does it: https://nand2mario.github.io/posts/2026/80386_barrel_shif... CPU-dependent behavior Posted Sep 29, 2026 16:00 UTC (Tue) by wtarreau (subscriber, #51152) [Link] (1 responses) CPU-dependent behavior Posted Oct 1, 2026 11:28 UTC (Thu) by khim (subscriber, #9252) [Link] I guess someone wanted to simplify something, but now we have the crazy story where the same operation either masks or doesn't mask depending on whether it's vectorized or not. For C it's not a problem: just make the developer suffer. It's all UB, anyway, thus it doesn't matter which operation is used. For Rust there are extra masking that you have to remove it by using intrinsics explicitly if your goal is to achieve as much performance as you want and you know what you are doing. CPU-dependent behavior Posted Sep 28, 2026 20:00 UTC (Mon) by ojeda (subscriber, #143370) [Link] Importantly, the key is that one should never rely on the wrapping nor the panicking. In other words, if one actually overflows, then it was not intentional, it is a bug, and one is supposed to fix the code. If one actually wants the panicking or the wrapping, then one should be explicit about it, rather than rely on that behavior. That way, the source code is not ambiguous, and thus can be analyzed well, while the program gets to keep perfectly well-defined semantics. No need to keep or add UB to a language just for that, as has sometimes been argued. So it is different than "normal" defined behavior. It is what nowadays one may call EB (Erroneous Behavior). I proposed naming this category in Rust back when Rust for Linux started and advocated it in C/C++/kernel discussions too, because I liked the idea from the integer operators in Rust. I presented it with that name in Kangrejos 2021, for instance. C++26 later adopted the concept under the same name, which is great. In Linux, on the Rust side, we use EB in certain places, e.g. we may document that a function is not supposed to be called in a certain way, but if it does get called badly, rather than to panic or to get the state corrupted, it will act in a defined way (possibly including reporting the error in the kernel log and returning some sort of fixed value). So one can see it as error handling for unwanted inputs, with the expectation that the callers should get fixed. CPU-dependent behavior Posted Sep 29, 2026 8:40 UTC (Tue) by ralfj (subscriber, #172874) [Link] (8 responses) The answer for Rust is that it depends on your compiler settings. By default, in debug builds it results in a panic which prints a stack-trace and exits. In release builds, it is guaranteed to wrap, which means it will evaluate to 0. The answer for Rust is that it depends on your compiler settings. By default, in debug builds it results in a panic which prints a stack-trace and exits. In release builds, it is guaranteed to wrap, which means it will evaluate to 0. No, it does not evaluate to 0. Did you actually try this before making your claims? https://play.rust-lang.org/?version=stable&mode=release&edition=2024&gist=34856eaababf1c620ea0bfb7f54b480c CPU-dependent behavior Posted Sep 29, 2026 13:46 UTC (Tue) by daroc (editor, #160859) [Link] (7 responses) Apparently that isn't the case, although a quick search does not seem to turn up any justification as to why. It does mean that replacing "* 2" with "<< 1" is not always a valid transformation, which seems a bit off. CPU-dependent behavior Posted Sep 29, 2026 19:03 UTC (Tue) by ralfj (subscriber, #172874) [Link] CPU-dependent behavior Posted Sep 30, 2026 6:48 UTC (Wed) by matthias (subscriber, #94967) [Link] (5 responses) CPU-dependent behavior Posted Sep 30, 2026 13:16 UTC (Wed) by daroc (editor, #160859) [Link] (3 responses) https://play.rust-lang.org/?version=stable&mode=relea... ... but the mechanism cannot possibly work like you describe, because of the last part of that example, which prints 1 instead of 0. If the operand were actually taken modulo the bit width, it would print 0. I think you are incorrect in the same way that I was at the beginning of this conversation, before ralfj corrected me. Unfortunately, it appears that Rust does not handle overflow for shifts in the same way that it handles overflow in other numerical operations. Personally, I think the current behavior ought to be considered a bug, and fixed to match what people clearly expect from Rust operators, but there would be a small performance penalty from doing so, and it would be a backward-incompatible change, so it might be a difficult case to make. CPU-dependent behavior Posted Sep 30, 2026 15:35 UTC (Wed) by NYKevin (subscriber, #129325) [Link] (1 responses) See [1] for an explanation of how Rust handles the << operator on signed integer types: > Panic-free bitwise shift-left; yields self << mask(rhs), where mask removes any high-order bits of rhs that would cause the shift to exceed the bitwidth of the type. > > Beware that, unlike most other wrapping_* methods on integers, this does not give the same result as doing the shift in infinite precision then truncating as needed. The behaviour matches what shift instructions do on many processors, and is what the << operator does when overflow checks are disabled, but numerically it’s weird. Consider, instead, using Self::unbounded_shl which has nicer behaviour. [1]: https://doc.rust-lang.org/std/primitive.i32.html#method.w... CPU-dependent behavior Posted Sep 30, 2026 18:52 UTC (Wed) by daroc (editor, #160859) [Link] CPU-dependent behavior Posted Sep 30, 2026 17:42 UTC (Wed) by ralfj (subscriber, #172874) [Link] For better or worse, shifts are not considered an arithmetic / numerical operation, but a logical / bitwise operation. I would have preferred the numeric version, but I am still surprised how many people here also expected the numeric version - I expected the LWN audience to be more low-level-minded. Sadly, not enough of us were present when this decision was made around 12 years ago... CPU-dependent behavior Posted Sep 30, 2026 17:34 UTC (Wed) by ralfj (subscriber, #172874) [Link] So the transformation is valid only if you assume that the original code never panics. CPU-dependent behavior Posted Sep 28, 2026 20:52 UTC (Mon) by azumanga (subscriber, #90158) [Link] I agree these things should be left cpu-specific, just not ‘oh you did this so now UB applies and I’m going to smash your whole program’. CPU-dependent behavior Posted Sep 29, 2026 4:00 UTC (Tue) by Aissen (subscriber, #59976) [Link] In Rust this is an overflow, and will panic like other int overflows if you build with integer overflow checking enabled (enabled by default in debug mode, and I'd argue most people should enable them in release mode). https://doc.rust-lang.org/reference/expressions/operator-... Additionally, Rust provides non-panicking variants for all integer types, with strictly defined semantics: https://doc.rust-lang.org/std/primitive.u64.html#method.w... https://doc.rust-lang.org/std/primitive.u64.html#method.c... https://doc.rust-lang.org/std/primitive.u64.html#method.u... https://doc.rust-lang.org/std/primitive.u64.html#method.s... or unsafe functions that do not panic: https://doc.rust-lang.org/std/primitive.u64.html#method.u... CPU-dependent behavior Posted Sep 29, 2026 13:00 UTC (Tue) by iabervon (subscriber, #722) [Link] (13 responses) In this particular case, I didn't find any way to get nasal demons out of gcc, although I only tried a little. (That is, I couldn't get gcc to produce non-zero but also eliminate code that would if it produced non-zero.) CPU-dependent behavior Posted Sep 29, 2026 23:24 UTC (Tue) by mathstuf (subscriber, #69389) [Link] (12 responses) In this particular case, I didn't find any way to get nasal demons out of gcc, although I only tried a little. (That is, I couldn't get gcc to produce non-zero but also eliminate code that would if it produced non-zero.) In this particular case, I didn't find any way to get nasal demons out of gcc, although I only tried a little. (That is, I couldn't get gcc to produce non-zero but also eliminate code that would if it produced non-zero.) What about comparing the shift in a way that is a tautology if it is always a valid shift value. Something like: int f(int s) { int n = 1 << s; if (s >= 32) sleep(60); return n; } UB says that the sleep can be optimized out. Does a compiler do so? CPU-dependent behavior Posted Sep 30, 2026 0:57 UTC (Wed) by iabervon (subscriber, #722) [Link] (11 responses) CPU-dependent behavior Posted Sep 30, 2026 1:17 UTC (Wed) by mathstuf (subscriber, #69389) [Link] (9 responses) I'd have expected it to be part of some bounded value optimization. For example, is the second conditional optimized out in: int a = x(); if (a < 10) f(); if (a < 5) g(); CPU-dependent behavior Posted Sep 30, 2026 1:54 UTC (Wed) by intelfx (subscriber, #130118) [Link] (5 responses) Unless f() is known not to return, I don't see what is there to optimize in this snippet? Perhaps you meant to write return f() or something in that direction? CPU-dependent behavior Posted Sep 30, 2026 10:37 UTC (Wed) by mathstuf (subscriber, #69389) [Link] (4 responses) If f doesn't return, the conditional is irrelevant. If it does return, it is tautological given a < 10 and doesn't need executed. Basically this can become if (a < 10) { f(); g(); }. CPU-dependent behavior Posted Sep 30, 2026 12:52 UTC (Wed) by dskoll (subscriber, #1630) [Link] (3 responses) No, that optimization breaks if a = 7. CPU-dependent behavior Posted Sep 30, 2026 14:57 UTC (Wed) by mathstuf (subscriber, #69389) [Link] (2 responses) Bah, should have flipped the 10 and 5 around, yes, sorry. That's what I get for late-night comment code :) . CPU-dependent behavior Posted Sep 30, 2026 15:10 UTC (Wed) by dskoll (subscriber, #1630) [Link] (1 responses) But even so, you can't completely optimize out the second test. Any such optimization breaks the case 5 < a < 10. At most, you could put the test for a < 5 inside the if part of a < 10, so that if a ≥ 10 you don't bother testing the a < 5 case. (Perhaps this is what you meant?) CPU-dependent behavior Posted Sep 30, 2026 15:45 UTC (Wed) by iabervon (subscriber, #722) [Link] CPU-dependent behavior Posted Sep 30, 2026 2:16 UTC (Wed) by iabervon (subscriber, #722) [Link] (2 responses) CPU-dependent behavior Posted Sep 30, 2026 10:39 UTC (Wed) by mathstuf (subscriber, #69389) [Link] (1 responses) True. I don't think using a bit shift amount as, say, an index into an array is common. Maybe some associated data with a specific bit in a bitfield? But then you're likely dealing with arrays and constants, not variables. CPU-dependent behavior Posted Sep 30, 2026 12:18 UTC (Wed) by iabervon (subscriber, #722) [Link] CPU-dependent behavior Posted Sep 30, 2026 15:03 UTC (Wed) by danielthompson (subscriber, #97243) [Link] int f(int s) { int n = 1 << s; if (s >= 32) sleep(n); return n; } There are some cases where sleep is passed a different argument (0) than is returned from f (1). I spotted this weirdness by inspection on RISC-V but was able to reproduce on x86-64 with Debian's gcc-14.2 compiler. CPU-dependent behavior Posted Sep 29, 2026 23:29 UTC (Tue) by Wol (subscriber, #4433) [Link] (3 responses) ??? What the writers of C *originally* intended aiui, is that what is now called "Undefined Behaviour" was meant to be "Defined Elsewhere". Just bring that back. At which point (for your example) you bring in a bunch of compiler flags such as "cpu-defined" (the default), "ones-complement", "twos-complement", "Z80" (for those who remember that bug) ... That simple "Defined Elsewhere" approach will probably get rid of nearly all UB at a stroke (it will take rather longer to actually get those definitions clarified and implemented :-) Cheers, Wol CPU-dependent behavior Posted Sep 30, 2026 6:23 UTC (Wed) by mb (subscriber, #50428) [Link] (2 responses) It is still like that. "Elsewhere" is the compiler defining "UB = Invalid program", which is a perfectly fine definition. CPU-dependent behavior Posted Sep 30, 2026 9:38 UTC (Wed) by taladar (subscriber, #68407) [Link] (1 responses) CPU-dependent behavior Posted Sep 30, 2026 15:37 UTC (Wed) by mb (subscriber, #50428) [Link] Division by zero, and other UB Posted Sep 28, 2026 17:33 UTC (Mon) by rrolls (subscriber, #151126) [Link] (13 responses) Pony does this, and I think provides a good justification: https://tutorial.ponylang.io/gotchas/divide-by-zero.html It does not matter that this definition is mathematically questionable. Defining a/0 to 0 is useful in _some_ circumstances (such as displaying an average, where one might output "N/A" if there are no samples, but outputting "0" is a very common choice), and the property of "all division results being defined" is useful in basically all circumstances (because it means you no longer need to worry about accidentally triggering UB!). The "downside" to defining the result as 0 is that suddenly compilers must add extra code, which means extra cycles consumed at runtime, to check if b is 0 prior to performing the division. However, this isn't actually a downside! People already have to write their own checks all over the place prior to performing a division, so a lot of code will already have those checks there. That means that the compiler would not emit any extra code, as it can see that by the time the division is reached, the check has already been done and b can't possibly be zero. It actually provides a benefit, because now compilers can also omit the check if the compiler can prove by some other means that b can't be zero. This technique can resolve other types of UB, too. Reading from an uninitialised variable or an out-of-bounds pointer could be defined to return 0. Calling a function pointer that is NULL could be defined to do nothing and return 0 (or the equivalent "zero value" for the return type). Writing out of bounds could be defined to do nothing. Similarly to the div/0 case above, compilers would suddenly need to produce extra code to check for these cases and thus slow down execution by default - but precisely because compilers would need to do that, they would _know_ when they are doing that, so can also emit a warning or error depending on your compiler options, to tell you that you're invoking these extra checks. If you know that a check isn't needed (for example if you know a pointer cannot be out-of-bounds), you can add an assert, and bam, the compiler now knows it can get rid of the check; your assert will perform the check if asserts are enabled, or you will get back your "undefined behavior" if asserts are disabled for performance. And if in your use case you don't need the performance, you could just not enable those warnings and let it silently add all the checks it wants, and you'll have guaranteed safe (even if slow) code. The obvious concern of "but zeroes will appear unexpectedly!" can then be addressed by getting into a habit of always assigning zero the meaning of "not special; default; disabled; resource not available; operation not known to be successful" or similar, and then any unexpected zero that shows up will just neatly send your program down its fallback or error path rather than doing anything dangerous. Division by zero, and other UB Posted Sep 28, 2026 18:41 UTC (Mon) by ballombe (subscriber, #9523) [Link] (7 responses) Division by zero, and other UB Posted Sep 28, 2026 19:12 UTC (Mon) by ballombe (subscriber, #9523) [Link] 35 years ago the C standard committee should have used its weight to encourage CPU with standardized behaviour, but instead they tried to beat fortran at its own game and completely failed. Division by zero, and other UB Posted Sep 28, 2026 19:15 UTC (Mon) by Cyberax (✭ supporter ✭, #52523) [Link] (5 responses) Division by zero, and other UB Posted Sep 28, 2026 20:19 UTC (Mon) by jengelh (subscriber, #33263) [Link] (1 responses) If only. Thanks to , it's not guaranteed. Division by zero, and other UB Posted Sep 28, 2026 21:30 UTC (Mon) by NYKevin (subscriber, #129325) [Link] Division by zero, and other UB Posted Sep 29, 2026 8:52 UTC (Tue) by malmedal (subscriber, #56172) [Link] (2 responses) Division by zero, and other UB Posted Sep 29, 2026 9:32 UTC (Tue) by pm215 (subscriber, #98099) [Link] (1 responses) The IEEE spec says the default is "set the flag and continue execution", but it also muddies the other waters by devolving various aspects of this to the individual programming language specs. Division by zero, and other UB Posted Sep 30, 2026 6:17 UTC (Wed) by malmedal (subscriber, #56172) [Link] Division by zero, and other UB Posted Sep 29, 2026 8:32 UTC (Tue) by alx.manpages (subscriber, #145117) [Link] (1 responses) Division by zero, and other UB Posted Sep 29, 2026 9:38 UTC (Tue) by ojeda (subscriber, #143370) [Link] It's preferable to entirely disallow division by zero when the operands are integer constant expressions, and UB at run-time, which analyzers can catch. It's preferable to entirely disallow division by zero when the operands are integer constant expressions, and UB at run-time, which analyzers can catch. For something like C, there is no need to use UB for analyzers to catch that -- they can (and already do) add checks like that just fine in practice, and if one wants to allow for that in the standard, one can introduce EB instead (if one is OK with a performance cost due to whatever defined behavior is chosen, of course). Either way, what is best is to provide the user with the ability to be as explicit as possible: if they need UB for a particular reason, let them ask for it explicitly; if they know zero shouldn't happen but they don't need the performance in that particular spot, then let them use EB; if they need particular handling on the zero case, then let them ergonomically do that; and so on. That is what also allows to easily read and understand (for both humans and tooling) what programs are supposed to do and what is unintentional. Division by zero, and other UB Posted Sep 29, 2026 11:26 UTC (Tue) by MortenSickel (subscriber, #3238) [Link] I cannot recall one single case when I have written something where a x/0 returning 0 would make sense. In some cases I have a denominator I know can be 0, then I have to check for it and handle the situation differently if it is 0 or not. In other cases, I may have a denominator that should never be 0, and if I feel brave enough, I do not test, or I may test just in case. What I definately do not want, is an accidential division by 0 returning something that afterwards looks like a reasonable answer, using that for some further calculations or checks and end up with a routine that seemed to run well but returns a nonsense value. (*) Quote from last Edinburgh fringe festival. Division by zero, and other UB Posted Sep 29, 2026 16:07 UTC (Tue) by wtarreau (subscriber, #51152) [Link] Division by zero, and other UB Posted Sep 30, 2026 15:02 UTC (Wed) by scott (subscriber, #581) [Link] Good news Posted Sep 29, 2026 1:56 UTC (Tue) by marcH (subscriber, #57642) [Link] (1 responses) It's easy to debate what programming languages everyone should use. Whereas this sort of work is barely visible and much more important - thank you! I don't want to choose between Rust and a safer C/C++: everyone should want _both_. They are not mutually exclusive and the more "security competitions", the merrier. Computers have been insecure and crashing for way too long and everything helps. Society will depend on C and C++ for at least a couple more generations. Good news - Subset of a superset Posted Oct 7, 2026 7:46 UTC (Wed) by swilmet (subscriber, #98424) [Link] https://herbsutter.com/2024/03/11/safety-in-context/ Nothing prevents from doing the same for the C language. But C evolves quite slowly compared to C++. Doubtful Posted Sep 29, 2026 7:47 UTC (Tue) by taladar (subscriber, #68407) [Link] (20 responses) The main benefit from using an older language like C (if there is any) is that the name identifies a certain language behaviour that existing C programmers know and existing code bases expect. So while it is of course possible to change all kinds of things and still call the result C it is questionable at best to do so if it costs you the compatibility with the existing code bases and the existing programmer's knowledge. At the same time though, it is orders of magnitude harder to ever get anywhere close to the benefits of a modern language like Rust that had the benefit of being able to start with a clean slate without having to consider backwards compatibility with existing code bases. Even if you could add enough features to the language that new code had most of the benefits, you would still have to live with having old code in your process (unless you throw away backwards compatibility completely in which case, why not call it something else) and that comes with a myriad of issues when enforcing any sort of safety or security guarantee. Doubtful Posted Sep 29, 2026 8:41 UTC (Tue) by alx.manpages (subscriber, #145117) [Link] (12 responses) I don't think it changes fundamentally. It's still the same thing, just changing some UB for compiler errors, and some other UB for defined behavior (the former is preferable). > > The main benefit from using an older language like C (if there is any) is that the name identifies a certain language behaviour that existing C programmers know and existing code bases expect. Old code that had defined behavior still works, and remains having the same meaning. The worst code using deep UB, it will probably now result in compiler errors; that's fine, since it never really worked. It can be fixed (and it can also be compiled under `-std=c89` if UB is indeed wanted). > At the same time though, it is orders of magnitude harder to ever get anywhere close to the benefits of a modern language like Rust that had the benefit of being able to start with a clean slate without having to consider backwards compatibility with existing code bases. According to Ojeda in Kernel Recipes, Rust already has applied breaking changes that silently change the behavior of some code that was valid and remains valid. That's something C has never done (IIRC) until very recently, where `auto` was changed to mean `__auto_type`. And it could be done with `auto` because no-one had seriously used it (there may be a few exceptions, but `auto` as automatic-storage duration was essentially unused). So, even with many decades of advantage, Rust has had to do these things it didn't need to. Doubtful Posted Sep 29, 2026 9:32 UTC (Tue) by ojeda (subscriber, #143370) [Link] (11 responses) According to Ojeda in Kernel Recipes, Rust already has applied breaking changes that silently change the behavior of some code that was valid and remains valid. According to Ojeda in Kernel Recipes, Rust already has applied breaking changes that silently change the behavior of some code that was valid and remains valid. To clarify: when migrating from one edition to another (particularly Rust 2021 to Rust 2024), i.e. it is an explicit change, not a silent one. The silent part was about backporting in Linux: if we tag a patch in Linux for backport, and the editions involved were to be such a pair, then the change would indeed be silent, because at the moment the process works in a way that makes it very likely nobody will notice the semantic change. And thus why I would like to have a tool that the stable kernel team can run to identify such changes so that patches can be flagged, similarly to how they are flagged when conflicts happen. I hope that clarifies. Doubtful Posted Sep 29, 2026 10:37 UTC (Tue) by alx.manpages (subscriber, #145117) [Link] (10 responses) To clarify: when migrating from one edition to another (particularly Rust 2021 to Rust 2024), To clarify: when migrating from one edition to another (particularly Rust 2021 to Rust 2024), Yup. i.e. it is an explicit change, not a silent one. i.e. it is an explicit change, not a silent one. That's more or less like changing the meaning of code from C17 to C23 (which has happened, with auto). In C, we'd call that a quiet change, because the one changing the language edition might not be aware of the code it is migrating, and thus of the implications of the change (essentially, the problem you face in the kernel with stable backports). Here's what the current C Charter says about silent changes: Avoid quiet changes Changes that alter the meaning of existing code cause problems. Breaking changes that require diagnostic messages are easily detected. Avoid silent changes that cause a working program to behave differently without requiring a diagnostic message. Where this principle is violated, informative notes should be added to the Standard. Avoid quiet changes Changes that alter the meaning of existing code cause problems. Breaking changes that require diagnostic messages are easily detected. Avoid silent changes that cause a working program to behave differently without requiring a diagnostic message. Where this principle is violated, informative notes should be added to the Standard. Avoid quiet changes Changes that alter the meaning of existing code cause problems. Breaking changes that require diagnostic messages are easily detected. Avoid silent changes that cause a working program to behave differently without requiring a diagnostic message. Where this principle is violated, informative notes should be added to the Standard. And here's what the C23 Charter said: Avoid “quiet changes.” Any change to widespread practice altering the meaning of existing code causes problems. Changes that cause code to be so ill-formed as to require diagnostic messages are at least easy to detect. As much as seemed possible, consistent with its other goals, the Committee has avoided changes that quietly alter one valid program to another with different semantics, that cause a working program to work differently without notice. In important places where this principle is violated, the Rationale points out a QUIET CHANGE. Avoid “quiet changes.” Any change to widespread practice altering the meaning of existing code causes problems. Changes that cause code to be so ill-formed as to require diagnostic messages are at least easy to detect. As much as seemed possible, consistent with its other goals, the Committee has avoided changes that quietly alter one valid program to another with different semantics, that cause a working program to work differently without notice. In important places where this principle is violated, the Rationale points out a QUIET CHANGE. Avoid “quiet changes.” Any change to widespread practice altering the meaning of existing code causes problems. Changes that cause code to be so ill-formed as to require diagnostic messages are at least easy to detect. As much as seemed possible, consistent with its other goals, the Committee has avoided changes that quietly alter one valid program to another with different semantics, that cause a working program to work differently without notice. In important places where this principle is violated, the Rationale points out a QUIET CHANGE. Which we've respected quite much, precisely because backporting issues are very problematic. The exception is below: auto x = 0.0; The meaning of the above in C89 was int x = 0.0;, and in C23 it means double x = 0.0;. Hopefully, nobody will write code like that and backport it; at least that's what the committee hoped. I'm dubious of this precise change, precisely because it adds some unnecessary risk. On the other hand, there's the risk reduction in that C++ code ported to C would now mean the same, so maybe it was a good change. Since that code is weird in the first place, and avoided by C programmers in general, I'm not too worried. I don't see any 'QUIET CHANGE' notes in C23, which I'll report as a bug in C23. Doubtful Posted Sep 29, 2026 12:01 UTC (Tue) by khim (subscriber, #9252) [Link] (3 responses) > In C, we'd call that a quiet change, because the one changing the language edition might not be aware of the code it is migrating, and thus of the implications of the change (essentially, the problem you face in the kernel with stable backports). Precisely — but that's because new features introduced in new C standards are not accessible for the code that's compiled for old C standard. You couldn't compile your code as C17 code and use _BitInt types, e.g. Also: C offer no way to compile some library as C17 code (including things like macros) and yet use it in C23 code — while Rust remembers which edition was in use when macro was defined to permit precisely that mix. It's expected and perfectly normal to combine code for different Rust editions in one program. This means that you may want to upgrade to later C standard to get useful features and then hit these “unexpected changes”. But in Rust that's not the case: you have access to all new features in all editions (except when new feature requires new syntax not supported by old edition). You only upgrade to a new edition specifically to opt in into these “silent behavior changes”. It's a bit silly to call behavior change “silent” when it's well documented and, more importantly, the whole reason new edition even exist! Doubtful Posted Sep 29, 2026 14:24 UTC (Tue) by alx.manpages (subscriber, #145117) [Link] (2 responses) Precisely — but that's because new features introduced in new C standards are not accessible for the code that's compiled for old C standard. You couldn't compile your code as C17 code and use _BitInt types, e.g. Precisely — but that's because new features introduced in new C standards are not accessible for the code that's compiled for old C standard. You couldn't compile your code as C17 code and use _BitInt types, e.g. Compilers often implement new features that are backwards compatible, even in standard mode. So, actually, you can, and it seems very similar to what you say of Rust. alx@debian:~/tmp$ cat bi.c int main(void) { _BitInt(8) i = 42; return 42; } alx@debian:~/tmp$ gcc -std=c89 bi.c alx@debian:~/tmp$ ./a.out; echo $? 42 Also: C offer no way to compile some library as C17 code (including things like macros) and yet use it in C23 code — while Rust remembers which edition was in use when macro was defined to permit precisely that mix. It's expected and perfectly normal to combine code for different Rust editions in one program. Also: C offer no way to compile some library as C17 code (including things like macros) and yet use it in C23 code — while Rust remembers which edition was in use when macro was defined to permit precisely that mix. It's expected and perfectly normal to combine code for different Rust editions in one program. Some experimental compilers do provide such a feature (IIRC). It's not something common or even desirable, but it exists. You only upgrade to a new edition specifically to opt in into these “silent behavior changes”. You only upgrade to a new edition specifically to opt in into these “silent behavior changes”. Hummmm, sounds like a huge problem in 2099, when there might be dozens of editions with slightly different behavior. If one doesn't update code unless a quiet change is needed, then most code might stay on old editions, and a few lines of code might use wildly different editions. Doubtful Posted Sep 29, 2026 23:13 UTC (Tue) by mathstuf (subscriber, #69389) [Link] Hummmm, sounds like a huge problem in 2099, when there might be dozens of editions with slightly different behavior. If one doesn't update code unless a quiet change is needed, then most code might stay on old editions, and a few lines of code might use wildly different editions. Hummmm, sounds like a huge problem in 2099, when there might be dozens of editions with slightly different behavior. If one doesn't update code unless a quiet change is needed, then most code might stay on old editions, and a few lines of code might use wildly different editions. There's no issue with that. It's not like Rust 2015 is going anywhere. It might be easier to think of newer editions allowing new spellings of existing code. Everything still describes the same underlying semantics, but an edition change can do things like: Offer new impls of existing types: for example, impl IntoIterator for [T; N] is new in Rust 2021 so that [1, 2].into_iter() now works (previously only .iter() worked which gave references to the items, not the items themselves). Rust 2015 and 2018 can't spell that directly, but Rust 2021 can pass [1, 2] to a Rust 2015 function requiring IntoIterator and it's all OK. Changing how use works. Rust 2015 mounted all of the crate's modules and dependencies at the "root" (e.g., use mylocalmod;). Rust 2018 puts the crate's own modules under a crate:: namespace instead (i.e., use crate::mylocalmod;). Since any given crate is a single edition, how dependencies and internal modules are loaded can be migrated crate-by-crate. Rust 2018 introduced the async and await keywords. If a Rust 2015 API had an async member, Rust 2018 can still use it by spelling it as r#async to "remove" the keyword-ness of it. There is also the k# prefix to force the keyword-ness powers of the token, but I'm not aware of any non-contorted use cases at the moment. For many of these things, cargo fix can help do the migration. I'm not aware of anyone blindly bumping the edition for their crate and not at least testing that things still work. During an edition bump, one usually kickstarts the process with cargo fix and/or cargo clippy as guidance. Doubtful Posted Sep 30, 2026 6:07 UTC (Wed) by mb (subscriber, #50428) [Link] Why would that be a huge problem? There are things that Editions cannot change. Like the public APIs of crates are interpreted in different incompatible ways by different editions. All editions are supposed to interact with each other. In the same program. There are currently many libraries using older editions out there and it's no problem at all. The only "problem" is that you end up with two dozen of actively supported edition implementations in the compiler in 100 years. So compiler complexity rises a bit. Up until now the actual differences between editions are not that big. Because they only effect the things that cannot be changed in a backwards compatible way. Editions are also not frozen in time. Old editions constantly receive all new features, if these features are backwards compatible to the edition. Doubtful Posted Sep 29, 2026 16:29 UTC (Tue) by ojeda (subscriber, #143370) [Link] Yes, I just wanted to clarify that it is not a random change that happened out of nowhere, because someone reading your previous message may have thought that without having the extra context. In Rust, editions are meant to introduce backwards-incompatible changes, and there is tooling to support the migrations. From the Rust edition guide: “Editions” are Rust’s way of introducing changes into the language that would not otherwise be backwards compatible. “Editions” are Rust’s way of introducing changes into the language that would not otherwise be backwards compatible. So, in Rust, someone migrating to the next edition generally knows that it takes extra steps and that they are supposed to run the tooling on their project, to read the guide, to clean certain lints in advance, etc. But, yes, to be clear, I am not thrilled about subtle semantics changes, even across an edition boundary, for the reasons I mentioned in the talk among others. And, yeah, I generally agree with those C Charter principles -- I was one of the authors, after all. Doubtful Posted Sep 29, 2026 19:11 UTC (Tue) by ralfj (subscriber, #172874) [Link] (4 responses) In Rust, the author of a crate decides on the edition of the crate. This change is aided by migration tools that adjust the code in a way that it retains its original meaning. It's not like in C where the "-std" flag might be set by someone else than the author of the code. The issue arises because of the very special needs of the kernel where it maintains both the old-edition code and the new-edition code after the migration, and furthermore moves patches between old and new versions. For almost all Rust code out there, it's not developed like this and there is no issue with semantics over edition migrations. Doubtful Posted Sep 29, 2026 22:30 UTC (Tue) by ojeda (subscriber, #143370) [Link] (3 responses) The issue arises because of the very special needs of the kernel where it maintains both the old-edition code and the new-edition code after the migration, and furthermore moves patches between old and new versions. The issue arises because of the very special needs of the kernel where it maintains both the old-edition code and the new-edition code after the migration, and furthermore moves patches between old and new versions. To clarify: we don't have the problem at the moment, i.e. I am holding off the move to the new edition precisely to avoid the issue. For almost all Rust code out there, it's not developed like this and there is no issue with semantics over edition migrations. For almost all Rust code out there, it's not developed like this and there is no issue with semantics over edition migrations. Yeah, it is not common, but as Rust gets more popular, I can imagine projects and companies out there that maintain LTS branches of a product for long enough hitting the problem sooner or later. Doubtful Posted Sep 30, 2026 15:43 UTC (Wed) by NYKevin (subscriber, #129325) [Link] (2 responses) Doubtful Posted Sep 30, 2026 15:49 UTC (Wed) by mb (subscriber, #50428) [Link] cargo fix can do that (after applying the patch). Of course it probably won't catch all problems. Doubtful Posted Sep 30, 2026 16:17 UTC (Wed) by ojeda (subscriber, #143370) [Link] It would be from new edition to old edition, i.e. the other direction, but yeah. We don't actually need the conversion (although it would be nice), "just" the detection. Doubtful Posted Sep 29, 2026 18:33 UTC (Tue) by wahern (subscriber, #37304) [Link] [1] e.g., "Proposal for a Friendly Dialect of C", https://blog.regehr.org/archives/1180 Perfection is not possible, but incremental improvement is Posted Oct 1, 2026 9:36 UTC (Thu) by joib (subscriber, #8541) [Link] (5 responses) However, I think it's wrong to say C cannot be improved. Languages evolve, or they die. Heck, even COBOL and Fortran are still evolving. C is much too widely used for it to be ready for the museum anytime soon, no matter what you or I might think of it's deficiencies compared to Rust or other more modern languages. Thus, I think there is scope for improving C. Yes, it will likely never reach Rust level of safety, but that doesn't mean incremental improvement isn't possible. And given the huge amount of C code out there that isn't being rewritten in Rust anytime soon, even incremental improvements matter. The current work in narrowing down the scope of UB seems exactly like that sort of work. I believe a lot of this current work done on the C standard is not about outlawing currently allowed code, but about requiring certain cases of UB to be either diagnosed as an error, or then behaving in a defined manner. Perfection is not possible, but incremental improvement is Posted Oct 2, 2026 7:43 UTC (Fri) by taladar (subscriber, #68407) [Link] Perfection is not possible, but incremental improvement is Posted Oct 2, 2026 18:10 UTC (Fri) by Cyberax (✭ supporter ✭, #52523) [Link] (3 responses) People like to talk about how Fortran is still alive, but it's not. It's dead. Approximately nobody is writing large amounts of new serious code in Fortran. Not even in areas like numeric simulations. It'll be around for a while because there are optimized libraries that nobody dares to touch, and they just tend to be passed down as holy artifacts. But there's still new code written in COBOL, mostly for old systems that just work. Perfection is not possible, but incremental improvement is Posted Oct 2, 2026 19:34 UTC (Fri) by mathstuf (subscriber, #69389) [Link] But there's still new code written in COBOL, mostly for old systems that just work. But there's still new code written in COBOL, mostly for old systems that just work. There's even standardization work with a new edition due in the next year or two. I think Fortran standardization still meets and has activity, so there's certainly still interest around there too. But I'd not be surprised if the only effects noticed in the FOSS world are flang, gfortran, and gcobol improvements. Perfection is not possible, but incremental improvement is Posted Oct 2, 2026 21:44 UTC (Fri) by joib (subscriber, #8541) [Link] Just in my own former niche, I could rattle off at least half a dozen actively developed simulation applications, both open source and proprietary, where Fortran is the main implementation language. And this would include the most commonly used ones in the field. But this is very much depending on which community you're part of. Most scientists are not that interested in programming languages for their own sake, they just pick whatever their colleagues are using. Perfection is not possible, but incremental improvement is Posted Oct 3, 2026 3:34 UTC (Sat) by Klaasjan (subscriber, #4951) [Link] The meaning of "undefined behaviour." Posted Sep 29, 2026 8:44 UTC (Tue) by rweikusat2 (subscriber, #117920) [Link] (13 responses) On a more practical note, a program has undefined behaviour the moment it calls any function not defined by the C standard. For instance, any POSIX function that's not part of it. But that's not an issue because while the C standard doesn't define any behaviour for such a case, something else, like POSIX, very well might. Undefined by the C standard means neither wrong nor bizarre, just not defined by this particular document. The meaning of "undefined behaviour." Posted Sep 29, 2026 10:07 UTC (Tue) by elaforma (subscriber, #165356) [Link] (12 responses) Comparing this to features/topics outside the scope of the standard is, IMHO, semantic bickering at best and actively misleading at worst. The meaning of "undefined behaviour." Posted Sep 29, 2026 12:50 UTC (Tue) by rweikusat2 (subscriber, #117920) [Link] (11 responses) behaviour, upon use of a nonportable or erroneous program construct or of erroneous data, for which this document imposes no requirements And "no requirements" obviously includes behaving in "a documented manner characteristic of the environment", as explicitly stated in the next sentence, eg, gcc -fwrapv. The meaning of "undefined behaviour." Posted Sep 29, 2026 22:57 UTC (Tue) by Cyberax (✭ supporter ✭, #52523) [Link] (10 responses) Like the classic: "c = c++ + c++"? The meaning of "undefined behaviour." Posted Sep 30, 2026 9:31 UTC (Wed) by rweikusat2 (subscriber, #117920) [Link] (9 responses) The meaning of "undefined behaviour." Posted Sep 30, 2026 17:16 UTC (Wed) by Cyberax (✭ supporter ✭, #52523) [Link] (8 responses) And there are plenty more. C is just not well-specified. The meaning of "undefined behaviour." Posted Oct 1, 2026 12:25 UTC (Thu) by rweikusat2 (subscriber, #117920) [Link] (7 responses) #include int main(void) { char a[] = "vitzliputzli"; void *p; p = a; ++p; printf("%c\n", *(char *)p); return 0; } has undefined behaviour insofar the C standard is concerned but with gcc, the behaviour is perfectly well-defined and it will print i. The meaning of "undefined behaviour." Posted Oct 2, 2026 0:43 UTC (Fri) by Cyberax (✭ supporter ✭, #52523) [Link] (6 responses) There is nothing whatsoever preventing the C language from rigorously specifying the order of operations, like Java or many other languages did. The meaning of "undefined behaviour." Posted Oct 2, 2026 8:43 UTC (Fri) by rweikusat2 (subscriber, #117920) [Link] (5 responses) c = c++ + c++; is an expression statement with no sequence point in it. Hence, the behaviour is undefined because Between the previous and next sequence point an object shall have its stored value modified at most once by the evaluation of an expression. [6.5|2] My example has undefined behaviour because For addition, [...] one operand shall be a pointer to an object type and the other shall have integer type. [6.5.6|2] and void * is a pointer to an incomplete type and not to an object type. In both cases, a shall-constraint is violated. My point was that undefined behaviour according to C isn't necessarily a program error as something other than the C standard may well define this behaviour. I don't quite understand what you're up to. One can conjecture that, by the time the C standard was first created, different implementations already treated such expressions in different ways. The meaning of "undefined behaviour." Posted Oct 3, 2026 3:25 UTC (Sat) by Klaasjan (subscriber, #4951) [Link] (4 responses) Implementation defined versus undefined behaviour Posted Oct 5, 2026 9:28 UTC (Mon) by farnz (subscriber, #17727) [Link] (3 responses) The argument that keeps getting made (that I personally disagree with) is that the standard might as well make things undefined behaviour instead of unspecified or implementation defined behaviour because all three of those are the same from the perspective of a strictly conforming program, unless you can put restrictions on the behaviour that allow a strictly conforming program to make use of this. For example, defining the size of unsigned as "implementation defined, must be at least 16 bits" is considered OK, because a strictly conforming program can be written such that it's fine with unsigned being 16 bits or larger. Leaving it as simply "implementation defined" would not be under this argument, because then it could be any size, and a strictly conforming program would not be allowed to make any assumptions about the size of unsigned. This argument rests on the idea that as a Quality of Implementation decision, a compiler can choose to define anything that's already undefined behaviour; just because a strictly conforming compiler sees UB when it sees c = c++ + c++ doesn't mean that GCC (for example) couldn't choose to define this as equivalent to c = c + c + 4; (which would surprise a naïve reader, since you might expect it to be c = c + (c+1); with the results of the ++ discarded). Implementation defined versus undefined behaviour Posted Oct 5, 2026 11:17 UTC (Mon) by mathstuf (subscriber, #69389) [Link] (2 responses) just because a strictly conforming compiler sees UB when it sees c = c++ + c++ doesn't mean that GCC (for example) couldn't choose to define this as equivalent to c = c + c + 4; (which would surprise a naïve reader, since you might expect it to be c = c + (c+1); with the results of the ++ discarded). just because a strictly conforming compiler sees UB when it sees c = c++ + c++ doesn't mean that GCC (for example) couldn't choose to define this as equivalent to c = c + c + 4; (which would surprise a naïve reader, since you might expect it to be c = c + (c+1); with the results of the ++ discarded). I think you need more than that. The order of operations need to be explicitly defined because you have: *square_matrix_iter = (*(square_matrix_iter++)) * (*(square_matrix_iter++)); where the order of operations is visible behavior. So the order of the increments, the dereference operators, and the assignment all need to be well-defined in some documentation. Then the optimizers need updated to do this. I imagine that there are some expression rewriters that assume that subexpressions cannot change during an evaluation based on the order of operations. Rust might actually have more trouble due to which subexpression panics on overflow being observable. Consider: let sum = arr[2] + arr[1] + arr[4] + arr[3]; I'm not sure if Rust can reorder this to arr.iter().copied().sum() (with possible vectorization open after that), in general, because an overflow can (theoretically) detect which expression overflowed. Floating point may also change results even without overflow based on the order of operations due to precision loss when values differ by a large magnitude. Implementation defined versus undefined behaviour Posted Oct 5, 2026 11:39 UTC (Mon) by mb (subscriber, #50428) [Link] I'm not sure if Rust can I'm not sure if Rust can The ordering only affects the bounds check, not the calculation itself. The calculation is done with SIMD and doesn't change with different orderings. https://play.rust-lang.org/?version=stable&mode=release&edition=2024&gist=b3d14c45a15957e87e53b9632d8205c0 That also means let sum = arr[4] + arr[3] + arr[2] + arr[1]; is the most efficient from a bounds check perspective as only one check for 4 is generated. Implementation defined versus undefined behaviour Posted Oct 5, 2026 12:18 UTC (Mon) by farnz (subscriber, #17727) [Link] You misunderstand my point - GCC (as an implementation) can take the UB in c = c++ + c++; and define it as having any meaning that GCC wants it to have, without the standard being updated to make it no longer UB. It could even be defined as having some form of non-deterministic behaviour if that suits GCC's optimizer better; you could say that the execution of i++ is defined by the implementation as starting the store when the increment operator is executed, and then putting a wait for the store to complete at the ;. That would make *square_matrix_iter = (*(square_matrix_iter++)) * (*(square_matrix_iter++)); have up to 4 different meanings, depending on when the increments actually completed as compared to the loads. Some of these defined variants are less useful than others, but the argument remains that it might as well be UB in the ISO specification, because an implementation can define it downstream of the ISO specification. The Rust case is different because Rust defines the behaviour there - panic in debug builds, overflow and wrap in release builds. n3957 - Ghosts and Demons: Undefined Behavior in C2Y (Status 2026-08-23) Posted Sep 29, 2026 8:48 UTC (Tue) by alx.manpages (subscriber, #145117) [Link] (2 responses) There's a mistake in UB 2, which was defined recently, but the definition is bogus, and the committee is discussing what to do with it. Hopefully, we'll transform that into a compiler error. Others are pushing to define it differently, although that's wrong (I have a draft of a committee paper for that). n3957 - Ghosts and Demons: Undefined Behavior in C2Y (Status 2026-08-23) Posted Sep 29, 2026 19:14 UTC (Tue) by ralfj (subscriber, #172874) [Link] (1 responses) n3957 - Ghosts and Demons: Undefined Behavior in C2Y (Status 2026-08-23) Posted Sep 30, 2026 10:04 UTC (Wed) by alx.manpages (subscriber, #145117) [Link] For example, there's this sentence in ISO C: > if an lvalue does not designate an object when it is evaluated, the behavior is undefined. But that sentence never triggers because either a contsraint is violated (and thus the code doesn't compile), or some earlier UB has been triggered. (n3882 has the specific details of this ghost; ) Time travel Posted Sep 29, 2026 20:36 UTC (Tue) by ojeda (subscriber, #143370) [Link] (11 responses) and some compilers have treated that way — but that compiler behavior was a bug. and some compilers have treated that way — but that compiler behavior was a bug. It seems my example from back then is still time traveling in the latest MSVC in Compiler Explorer -- for those who want to see it in action: https://godbolt.org/z/xxMvPErfT Clang seems to agree it is surprising, because it does introduce time travel only if one removes the volatile. (I would never have expected to see the line I suggested to induce time travel in MSVC appear in LWN years later!) Time travel Posted Sep 30, 2026 17:51 UTC (Wed) by ralfj (subscriber, #172874) [Link] (10 responses) Time travel Posted Oct 1, 2026 1:56 UTC (Thu) by NYKevin (subscriber, #129325) [Link] (9 responses) I'm more specifically worried about generics, of course. Nobody is going to manually construct such bizarre specializations, but they could appear as a result of chicanery like String being passed to a routine that accepts impl FromStr, and then Err is ! or Infallible. Then you can end up deleting a lot of monomorphized code because it tries to interact with a type that turns out to be uninhabited. You don't really have a choice about deleting that code, because the Err case has no object representation. There is no correct way to emit Err-handling code even if you wanted to (you could panic, I suppose, but how do you even detect the Err case if there's no discriminant?). Which reminds me rather a lot of the whole "deleting null pointer checks before inevitable dereference" problem. But, on the other hand, control can't get there without somebody constructing a ! instance first, and that's already UB. So maybe I'm getting spooked over nothing. Time travel Posted Oct 1, 2026 7:49 UTC (Thu) by taladar (subscriber, #68407) [Link] (7 responses) Time travel Posted Oct 2, 2026 11:13 UTC (Fri) by mathstuf (subscriber, #69389) [Link] Sounds like a field of ROP gadgets to me. Time travel Posted Oct 2, 2026 22:01 UTC (Fri) by NYKevin (subscriber, #129325) [Link] (5 responses) Option is, in English, "an Option that is always None." Because it is always None, the layout optimizer removes the enum discriminant, and now there is nothing we can inspect at runtime to figure out which variant is active. That doesn't cause a problem for safe Rust (and correctly-written unsafe Rust) because the None variant should always be active, and the compiler hard-codes it as such. In fact, with no enum discriminant and no payload, the resulting data structure takes up no memory at all. Any instructions that would read from or write to it have to be totally elided, because under the Rust memory model, a zero-byte read or write of an arbitrary non-null pointer is always a no-op (regardless of whether the pointer points into a valid memory allocation). If we insist that we have to emit something, we would then run into the problem that no widely-deployed architecture will let you do zero-byte loads or stores (except perhaps as fancy no-ops). Again, this does not cause problems for safe and correct-unsafe Rust, because None is a constant and we can always constant-fold it out of existence. TL;DR: It's not merely that emitting code for this case is slower or otherwise a bad idea. There is no code that we could possibly emit for this case, because the data we're trying to manipulate does not exist at runtime. Panic? Posted Oct 4, 2026 23:43 UTC (Sun) by gmatht (subscriber, #58961) [Link] (4 responses) Panic? Posted Oct 5, 2026 2:57 UTC (Mon) by mathstuf (subscriber, #69389) [Link] (2 responses) Isn't that the "can't the compiler just assert on UB occurring?" question C and C++ get all the time? Panic? Posted Oct 5, 2026 8:21 UTC (Mon) by gmatht (subscriber, #58961) [Link] (1 responses) Panic? Posted Oct 5, 2026 8:42 UTC (Mon) by mb (subscriber, #50428) [Link] That seems like reasonable behaviour for the case where the compiler can detect That seems like reasonable behaviour for the case where the compiler can detect Yes, but in your question was about panic on is_some() on an Option object. This type has a size of zero. It cannot store any information at runtime. It's always None. It's not possible to emit code for is_some(). Even if the compiler would emit a type with size=1 containing a discriminant being always 0. Why emit additional code checking discriminant != 0 that cannot occur? This would only add additional failure modes for unsafe Rust handling this type. Can't panic on `Option` Posted Oct 5, 2026 9:10 UTC (Mon) by farnz (subscriber, #17727) [Link] In this case, it can't even do that - Option is an Option whose only possible value is None - there isn't a Some(!) case at all. It's like saying the compiler should emit code to check if a 64-bit register contains a value greater than 265, and panic if the register contains that value - sure, the panic part is easy, but how do you implement the branch on AMD64? Time travel Posted Oct 1, 2026 8:42 UTC (Thu) by ralfj (subscriber, #172874) [Link] Exactly. That's why I would say there is no time-travel here. That said, UB can still be unintuitive. In particular, note that only time-travel around *observable behavior* is forbidden. The compiler can still reorder `let v = *x;` and `let u = *y;`, so if one of them is null that UB may seem to be time-traveling around the other memory access. This is allowed because (non-volatile) memory accesses are not considered "observable". The compiler will even optimize away the access entirely if the value is unused. So which memory accesses actually happen when and in which order is entirely up to the compiler, as long as in the end the program still behaves as-if it did all the operations you wrote in the program (under the usual non-UB assumption). Working link to the slides? Posted Sep 30, 2026 10:32 UTC (Wed) by alison (subscriber, #63752) [Link] (1 responses) Working link to the slides? Posted Sep 30, 2026 11:04 UTC (Wed) by mathstuf (subscriber, #69389) [Link] That's a 50x error of some kind (e.g., the reverse proxy is up but what actually serves the content is down). It usually means infrastructure is having a bad day somewhere.