Conversation

I’m a pointer width truther, you often don’t need 64 bits and it’s a wasteful default that makes poorer use of cache and your programs slower too

3
1
0

@hailey apparently you can compile a Rust program to a 32bit architecture and it's forward compatible to run on 64bit systems. I wonder if that makes a performance difference, or if e.g. the removal of SIMD support results in a slower program overall

1
0
0

@seanlinsley targeting straight up i686 isn’t what you want due to removal of simd like you say, also fewer and smaller registers. Check out x32 on Linux for the best of both worlds, it’s full x86_64 but with 32 bit only pointers

1
1
0

@hailey yowasp.org provides wasm versions of yosys and nextpnr that are comparable (nextpnr is more or less as fast) in perf to native versions built with -O3 because they effectively use the x32 ABI

1
0
0

@hailey i don't know about making it the default though. maybe if we dropped the whole .so/.dll situation and ran every library in its little 32-bit sandbox, communicating via low-overhead RPCs. but that has other overheads

remember Windows AWE? that was such a not fun time for everyone involved

1
0
0

@hailey @seanlinsley There is a x32 port of debian which is slowly dying due to lack of interest. If I remember correctly, most performance boost were seen with haskell but nearly negligible for most other programs.

1
0
0

@josch @seanlinsley on a ruby app I saw nearly half (650mb -> 350mb !) the memory use!

0
1
0

@hailey how would you build a system on x32 ABI that doesn't choke if you ask it to load an unusually large image? only way i can think of involves bringing back far pointers and i don't think that's a good tradeoff

3
0
0

@whitequark @hailey pls no segment registers

1
0
0

@gsuberland @hailey i guess every library knows its data segment load address and stuffs it into the high part of every pointer-containing register on load

you could use rip to save a register on x86 even, at the cost of having to share text and data segments (... or you could use (rip>>32)+1, which gives you free execute-only memory, ish)

1
0
0

@whitequark I mean apps that need to stay 64 bit can stay 64 bit. I think regardless of pointer size most apps would choke on such a large file anyway unless they’re explicitly designed to handle it, and if they’re designed to handle it they know they need 64 bit pointers

1
1
0

@whitequark @gsuberland segmentation gets such a bad wrap… just rebrand it ‘the actor model’ and programmers will think it’s the future

3
1
0

@hailey i wonder how using x32 with a separate copy of the libc, the GUI toolkit, Mesa, and all the other stuff you'd need to build twice compares to the savings you'd get from being more cache-efficient

1
0
0

@hailey @gsuberland 'the component model', in case of wasm :)

0
0
0

@whitequark @gsuberland for real though every library living in its own segment is an interesting idea, it might go some way to containing the blast if there’s a critical vuln in some library. especially if the OS restricts how/where segment registers can be loaded

2
1
0

@hailey @gsuberland i mean i know it's not literally what you said but that's probably the closest thing you can ship that's still portable to aarch64

1
0
0

@hailey @whitequark you could store information about what the pointer is allowed to do in the upper bits, and then... wait, hang on, we're just reinventing CHERI here

1
0
0

@gsuberland @hailey not fair to bring up my job when i'm on vacation :p

0
0
0

@hailey @gsuberland i'm not sure i'd use segment registers even if i was only using x86; you'd have to reload LDT on every jump in order to enact useful sandboxing, and i think the performance of that would be abysmal. but maybe i'm missing something?

1
0
0

@hailey i've spent the last ten minutes thinking about it and i don't think either qt or gtk would take very well to thunking :s

1
0
0

@hailey @gsuberland i guess you could have one big LDT for every library if you used a trusted compiler or verified all executable code to only contain appropriate far call instructions. that's kind of a tall order though. maybe it would still work out for opportunistic mitigation...

0
0
0

@astraluma @hailey tbh if all you did was to recompile every copy of Electron on an average system to use x32 you'd get most of the benefits of this already for the average user :s

0
0
0

@whitequark @hailey Arbitrary limits. The whole UI screeching to a swapping halt or the file manager crashing with OOM and bringing down the desktop environment is not a reasonable thing to happen when you encounter a 100000x100000 PNG file either.

The few apps that actually need to work with giant data can be 64-bit or work in patches from backing in the filesystem rather than memory. Things that just need to load giant images to *display* them can stream the decode and resample/reencode to a cached bounded-resolution jpeg during the loading process.

1
0
0

@dalias @hailey the specific problem i have with this approach is that it is anti-user. i have lived through it. the amount of applications that would use the equivalent of Microsoft AWE is "fuck all", and having your CAD software, your image editor, or even your text editor crash because you dared to have two gigabytes of something open is an experience I'm willing to pay in performance just so that I never have to live through it again

1
0
0

@whitequark @hailey Having it crash is anti-user, but effectively that's what happens when you don't have arbitrary limits.

Either you OOM and crash, or you allocate so much swap-backed memory that it takes 30 minutes to get back any resemblance of control over the system.

Just telling the user "this data is too big, you need to break it into workable-sized pieces or let the application break it into levels of detail where everything but what you're currently focused on is vastly-reduced-detail" is something I'd deem much more user-friendly.

But I agree there's room for disagreement here. Nobody's stopping you from having 64-bit apps. I'd just like to have things that also fit in 32-bit as long as you stick to reasonable data size, *and* the ability to handle huge data by rejecting it or working with reduced-detail rather than OOM-crashing or forcing me to upgrade.

1
0
0

@dalias @hailey we were discussing 64-bit being the wrong default, meaning: unless somebody specifically complained in a way where the complaint stuck, i should expect all of my software to arrive as the default of 32-bit binaries

i have lived in that world and i did not like it. and the alternative wasn't "allocate all of your swap and crunch the HDD for half a hour", the alternative was "use slightly more memory than the arbitrary and fragile 2GB limit would let you".

a CAD system letting you do that is a difference between hitting your tolerances and having to file down all the edges. sure, maybe you'll have to wait for a bit because this wasn't the anticipated workflow. but you did get a file you could then use. this is a real scenario from my actual life when i was starting out with CAD

1
0
0

@hailey i'd like x32 abi, but long mode gives more registers and i wanna keep them :P

1
0
0

@dalias @hailey i am however sad that the x32 ABI on Linux is seemingly completely dead because opt-in x32 is very useful and gives a performance boost to those compute-heavy applications where you can have a good idea of how big the working set is, or where the end user is sophisticated enough to know which version to run

1
0
0

@dalias @hailey i don't think you even need kernel support for it because you could use one of the... seccomp variants, i think?, to mangle the syscall parameters

but as far as i recall it got ripped out of every compiler already

2
0
0

@ariadne you can still use rex.r instructions in x32 code, no? am i missing something?

1
0
0

@whitequark you can, hence x32 ABI is fine. legacy x86 not so much.

0
0
0

@whitequark @dalias it’s still in gcc but it’s often unavailable due to not even being compiled into the kernel (arch) or gated behind a boot param (debian)

0
1
0

@whitequark @hailey I think you're thinking of the ilp32 aarch64 thing that was never finished. I have not heard anything about x32 removal from compilers.

1
0
0

@whitequark @hailey And I'm not sure if you can emulate x32 from userspace with seccomp. I think there are data structures where the copy between user and kernel space differs where you would need temp storage to emulate.

1
0
0

@dalias @hailey yeah you'd have to do that (but you do get more portability in exchange)

0
0
0

@hailey but then where would I get my ~3 free bits for pointer tagging? 😔

1
0
0

@geist is 2 bits good enough? Otherwise if your tagged objects are at least 8 bytes just 8 byte align them anyway :)

0
1
0
@hailey it's nice to have practically unlimited virtual address space though
0
0
0

@hailey

the 3 genders:
- PTR
- FARPTR
- REALLYFARPTR

0
1
0