C++ - Reddit
227 subscribers
48 photos
8 videos
1 file
25.4K links
Stay up-to-date with everything C++!
Content directly fetched from the subreddit just for you.

Join our group for discussions : @programminginc

Powered by : @r_channels
Download Telegram
muopt: a C++17 micro-library for argument parsing

Hello! I wanted to share a project that I've been working on for the last few days, a small argument parsing library called muopt. It started from a quite popular Rust crate called lexopt that I use in one of my CLI Rust projects. I wanted something similar to lexopt and couldn't find it for C++, so I decided to write my own.

muopt is more of a tokenizer/lexer than it is a parser. It doesn't "connect" a value to its option and just views them as an argument. muopt is intended to be minimal and bare-bones. Below is a really simple example of muopt in action.

#include "muopt/muopt.hpp"
#include <iostream>

int main(int argc, char argv) {
muopt::Parser parser(argc, argv);

while (std::optional<muopt::Arg> arg = parser.next()) {
if (arg->islong("help"))
std::cout << "usage: example [--input FILE]\n";
if (arg->is
long("input"))
std::cout << parser.argvalue().valueor("") << '\n';
}
}

The code above will match the arguments --help and -h. Both will print out the usage message. It will also match --input=Hello or --input Hello and print out the value. The README has a deeper explanation on how muopt works and the repo has more examples.

muopt is still very early; it's not even released yet. I'm still debating on how to take on certain things, for example the arg_value() return type and how errors should be handled. Any suggestions are welcomed!

GitHub: https://github.com/secona/muopt

https://redd.it/1v6342h
@r_cpp
Modern c++: operations on function pointers will break the constexpr function.

I'm on MSVC/c++23.

I used some trick to generate type id at compile time:

export template<typename... T> void typeid() {}
export using type
idt = void(*)();
//
type
idt oneid = typeid<int>;


I’ve found that whenever I try to cast a \`type\
id` value to an integer or simply perform a magnitude comparison within a `constexpr` function (even when the call chain is entirely resolved at compile time), the function ceases to be `constexpr`, or MSVC simply ignores the relevant code.

template<typename A, typename B>
constexpr auto someconstexprfunction()
{
typeidt a = typeid<A>;
type
idt b = typeid<B>;

//Total if block will be ignored, and the function will not be constexpr anymore.
if (a < b)
{
//discard
}
else
{
//discard
}

//cast the pointer to sizet then the function will not be constexpr anymore.
return (size
t)a;
}

How to fixed the case?

Or another type id solution at compile time will be nice(I have try so many implementation, none of them work correctly in compile time).

https://redd.it/1v6clnl
@r_cpp
gluq9Ol.gif.png
33.8 KB
Did you know you can use (dynamic) libraries inside clang-repl?
https://redd.it/1v6fdv0
@r_cpp
SUCO – Lightweight Distributed C/C++ Compiler Grid with intelligent SSD cache (alternative to Icecream/distcc)

Hi everyone,



I'd like to share my open-source project **SUCO** (SUper COmpiler Grid).



It's a lightweight, zero-config distributed C/C++ compilation and caching system for local networks. Designed as a simpler and faster alternative to Icecream/distcc and proprietary tools like IncrediBuild.



**Key features:**

\- Automatic worker discovery via UDP

\- SHA-256 based SSD cache → warm rebuilds can be up to \~4× faster

\- Windows client can cross-compile to Linux MinGW workers

\- Load-aware scheduling + automatic failover

\- Live web dashboard + Prometheus metrics

\- Easy CMake & IDE integration



Cold builds are roughly on par with Icecream, but cache hits skip the expensive compiler phases completely.



GitHub: https://github.com/MicBur/suco



I'd really appreciate any feedback, criticism or suggestions!

https://redd.it/1v7ce8q
@r_cpp
I built a C++ options execution and back testing engine with Black-Scholes. Looking for feedback on my concurrency model and memory layout.

Hi everyone,


I recently built a C++ options execution and backtesting engine utilizing Black-Scholes pricing, walk-forward optimization, and regime-based strategy selection. It's fully wired to execute paper/live trades via Alpaca.


I'm looking for harsh, constructive feedback on my architecture—specifically regarding my order ledger concurrency model and C++ memory layout.


Specific areas I'd love feedback on:
1. Concurrency in "Orderledger":
I opted for standard `std::mutex` and `std::atomic` flags rather than a lock-free queue for position state management. I'm prioritizing preventing "ghost fills" over sub-millisecond latency. Did I make the right trade-off here, and are there any glaring race conditions I missed?
2. Memory Layout:
I’m relying on standard containers (`std::vector`, `std::unordered_map`) but minimizing dynamic heap allocations on the critical path. Are there better patterns I should adopt to avoid OS allocator pauses during market hours?
3.
Black-Scholes Math Precision:
I used standard `double` precision and `std::erfc` rather than fast-math approximations to ensure numerical stability when calculating Greeks like Gamma on short-dated options.


Here is the repo: github.com/kdyn-ctrl/Nox


If you want to see the rationale behind my decisions, check out `docs/guides/DESIGN_THINKING.md` in the repository.
I also have journals of my work and guides if any other info is needed. I'm still quite novice so please be as harsh as you can I'm really trying to improve.


Thanks in advance for your time and feedback!

https://redd.it/1v7v4ev
@r_cpp
I built a 3D network router that "senses" attacks instead of just watching an alarm number

Built a 3D grid network simulator in C++. Packets move between nodes and reroute around traffic jams. On top of that, I built a controller that watches the overall behavior of the network — not just one alarm number — using 4 signals together: how full it is, how much is getting dropped, how fast it recovers after a spike, and how much room is left.

I compared it against a normal "alarm" style monitor — the kind most systems use, which only reacts once one number crosses a fixed line (here: more than 5% of traffic dropped, or the network more than 40% full)..

I threw 5 different fake attacks at both systems, across two network sizes, and added up the totals:

\\-Packets delivered:

my controller 1,091,830 vs normal monitor 613,472

(mine delivered 1.8x more)

\\-Average traffic lost, across ALL 5 attacks: mine (23% ) vs. normal monitor (28%)

(Quick note on those two percentages — some attacks are brutal and cause heavy loss no matter what, others are quiet and barely register. 23% and 28% are just the average across all of them combined. The real story is in the next part.)

Two of the five attacks — a slow quiet flood, and a sneaky targeted overload — never dropped more than 1% of traffic. That means the normal alarm-style monitor never noticed them at all, because it was only built to react once something gets loud, and these two stayed quiet on purpose.

My controller still caught both of them. Even though no single number spiked, the overall pattern still looked different from what normal traffic looks like — kind of like noticing something feels "off" about a room, even when nothing's visibly broken.


Code: https://github.com/Antriksh005/arena3d.

Been working on this for now some months , so Genuinely want criticism, especially on how I'm measuring "detection" — that part deserves scrutiny.

https://redd.it/1v7xg86
@r_cpp
New C++ Conference Videos Released This Month - July 2026 (Updated to Include Videos Released 2026-07-20 - 2026-07-26)

C++Now

2026-07-20 - 2026-07-26

Generic Programming for Multidimensional Arrays - The Boost.Multi Experiment to Integrate MD Arrays with STL Algorithms in the CPU and GPU - Alfredo A. Correa - [https://youtu.be/WriSoV4nu5E](https://youtu.be/WriSoV4nu5E)
Typed Linear Algebra - How to Not Crash on Mars - François Carouge - https://youtu.be/xZO7X8LH6Dg
From Template Metaprogramming to User Convenience - API Design Stories - Ruslan Arutyunyan - [https://youtu.be/fPo\_Tff-L5Y](https://youtu.be/fPo_Tff-L5Y)

2026-07-13 - 2026-07-19

Keynote: Benchmarking - It's About Time - Matt Godbolt - https://youtu.be/EU\_nQh8wg5A
The Morning Briefing - C++ Concurrency Before the Hardware Reckoning - Fedor Pikus - [https://youtu.be/WtChBezurj8](https://youtu.be/WtChBezurj8)
Lock-free Programming is Dead - Long Live Lock-free Programming! - Fedor G Pikus - https://youtu.be/UdKqfQ3a\_sY

2026-07-06- 2026-07-12

Keynote: Reflection Is Only Half the Story - Barry Revzin - [https://youtu.be/DZTkT1Cq\_aY](https://youtu.be/DZTkT1Cq_aY)
Keynote: Multidimensional Parallel Standard C++ - Mark Hoemmen - https://youtu.be/VAwW\_s1uEHY

C++Online

2026-07-20 - 2026-07-26

When One Red Pill Is Not Enough - Compile-Time Optimization Through Dynamic Programming - Andrew Drakeford - [https://youtu.be/zUKqoBq6Sg4](https://youtu.be/zUKqoBq6Sg4)
Coroutines and C++ - Async Without The Pain? - Tamas Kovacs - https://youtu.be/4keXOkbr0UY

2026-07-13 - 2026-07-19

C++ Contracts - A Meaningfully Viable Product, Part II - Andrei Zissu - [https://youtu.be/1VUqOx6PCMU](https://youtu.be/1VUqOx6PCMU)
Sanitize it Before You Ship IT - Vishnu Nath - https://youtu.be/jzcGATR78Mk

2026-07-06 - 2026-07-12

Time to Introspect - A Beginner's Guide to Practical Reflection - Sarthak Sehgal - [https://youtu.be/9stn1o149pw](https://youtu.be/9stn1o149pw)
Refactoring Towards Structured Concurrency - Roi Barkan - https://youtu.be/6502xFEreI8

2026-06-29 - 2026-07-05

Why std::vector Can't Save You (And What to Use Next) - Kevin Carpenter - [https://youtu.be/78fYPix0mN4](https://youtu.be/78fYPix0mN4)
Modern C++ for Embedded Systems - From Fundamentals To Real-Time Solutions - Rutvij Karkhanis - https://youtu.be/XxeqHRDhHkU

ADC

2026-07-20 - 2026-07-26

Enumerate and Extract Audio Buffers When Debugging C++ Applications - Henning Meyer - [https://youtu.be/UHV4U\_ivm\_8](https://youtu.be/UHV4U_ivm_8)
A Sine By Any Other Language - David Su - https://youtu.be/5yEd1q\_\_cqo
When Code Writes Back: Making AI Coding Agents Actually Work - Tobias Baumbach - [https://youtu.be/K04ehohSdXc](https://youtu.be/K04ehohSdXc)
Mind the Spike - Benchmarking for Worst-Case Execution Time in Realtime Code - Christian Luther - https://youtu.be/7RrOjl996WQ

2026-07-13 - 2026-07-19

Building Inclusive Audio Tools - Accessibility with ARIA, WCAG, and Real-World Projects - Samuel John Prouse & David Shervill - [https://youtu.be/O5xX9a7P-SU](https://youtu.be/O5xX9a7P-SU)
PSD to DAW - Building a Pixel-Perfect UI Pipeline - Bence Kovács - https://youtu.be/hebLkAR5X3I
Analog Filters for Realtime Audio - George Gkountouras - [https://youtu.be/NLt0NqUtNLo](https://youtu.be/NLt0NqUtNLo)
A History of FLAC - The Free Lossless Audio Codec - Josh Coalson - https://youtu.be/tBb1STRW56s

2026-07-06 - 2026-07-12

Workshop: Audio Plugin DSP in Practice - Jan Wilczek & Linus Corneliusson - [https://youtu.be/Atc0GRWoolI](https://youtu.be/Atc0GRWoolI)
An Open Toolkit for Real-Time Audio Descriptors - Valerio Orlandini -
https://youtu.be/HKlnn0hd8J0
Bugs I’ve Seen in the Wild - From Confusion to Amazement - Olivier Petit - [https://youtu.be/LBWtb\_uXt0I](https://youtu.be/LBWtb_uXt0I)
Real-Time Audio in Python: Introducing the asmu Package - Felix Huber - https://youtu.be/X2vr81CJ934

2026-06-29 - 2026-07-05

Beyond iLok: Advanced Code Protection and Cryptography for the Next Generation - Protecting the Next Generation of Applications, Plug-ins, and AI Models - Neal Michie, Ryan Wardell & Bob Brown - [https://youtu.be/dbbK\_ry2cgo](https://youtu.be/dbbK_ry2cgo)
Database Synchronisation for Audio Plugins, Part Two - Here's One I Made Earlier - Adam Wilson - https://youtu.be/wJCy2G969ro
Perfect Oscillators in Less Than One Clock Cycle - Angus Hewlett - [https://youtu.be/Ssq0a-YdamM](https://youtu.be/Ssq0a-YdamM)
Driving Chaos - Virtual Analog Modelling of a Chaotic Circuit with Wave Digital Filters - Francisco Bernardo - https://youtu.be/PnEZNqyKlIw

Boost Documentary

There is also a teaser trailer for a new documentary on the history of the Boost C++ library https://www.youtube.com/watch?v=87jvuDbnwqQ which will have its first showing at CppCon this year

https://redd.it/1v83trq
@r_cpp
AVX-512 Intrinsics Optimization: Achieving ~7% speedup over libpopcnt on Sapphire Rapids

Benchmarking random 64B cache-line jumps on Intel Xeon Sapphire Rapids showed the custom kernel achieving 7.09 ns/line, compared to libpopcnt (7.58 ns) and std::popcount (46.10 ns).

While contiguous memory performance ties libpopcnt at 0.47 ns/word, the performance gain on this gather hot-path comes from maintaining accumulator states inside ZMM registers across steps and applying 2-accumulator unrolling to break dependency chains.

Target is v40. Looking for insights on whether software prefetching (_mm_prefetch) provides measurable benefits for this gather pattern or if latency is memory-bound.

https://redd.it/1v86xcs
@r_cpp
Ray Tracing in One Weekend in Metal

Hi! I recently implemented Ray Tracing in One Weekend in Metal using metal-cpp. Resources for metal-cpp specifically are pretty sparse, so I figured I'd share in case it helps anyone getting started with it or with graphics programming in general. Hope it helps!

Repo: https://github.com/jinhgkim/Path-Tracer

https://redd.it/1v8kbu7
@r_cpp
Constraints of mdspan policy layout_stride

Wtith great support of Mark Hoemmen and Christian Robert Trott (two authors of std::mdspan), I just finished the mdspan chapter of "C++23 - The Complete Guide" (https://www.cppstd23.com/). I learned something I was not aware of and what is not obvious:

If you specify a layout\_stride mapping, there are several constraints you have to take into account. Otherwise, the resulting code has undefined behavior.

The constaintes are:

* Each layout mapping must be ***unique***. This means that each underlying element may only be reached by one combination of indices. In other words: subsets of this layout may not overlap and iterating over all elements may not visit any element twice.
* Only ***positive offsets*** from one stride to the other are allowed. This, for example, means that you cannot use this layout for direct reverse iterations.
* There must be one way to perform a multi-dimensional iteration so that the ***resulting offsets*** to the underlying memory are **ascending.**

For example:

* [https://www.cppstd23.com/code/mdspan/mdspanstride2.cpp.html](https://www.cppstd23.com/code/mdspan/mdspanstride2.cpp.html) is valid.
* [https://www.cppstd23.com/code/mdspan/mdspanstrideUB4.cpp.html](https://www.cppstd23.com/code/mdspan/mdspanstrideUB4.cpp.html) has undefined behavior.

Hope this helps.

https://redd.it/1v8zynw
@r_cpp
Latest News From Upcoming C++ Conferences (2026-07-28)

**TICKETS AVAILABLE TO PURCHASE**

The following conferences currently have tickets available to purchase

* **CppCon (12th – 18th September)** – You can buy standard tickets until August 29th at [https://cppcon.org/registration/](https://cppcon.org/registration/)
* **C++ Under The Sea** **(14th – 16th October)** – You can buy early bird tickets at [https://sales.ticketing.cm.com/cppunderthesea2026/](https://sales.ticketing.cm.com/cppunderthesea2026/)
* **ADC – (9th – 11th November)** – Tickets for ADC can now be purchased at [https://ti.to/audio-developer-conference/adc-bristol-2026](https://ti.to/audio-developer-conference/adc-bristol-2026)
* **Meeting C++ (26th – 28th November)** – You can buy early bird tickets at [https://meetingcpp.com/2026/](https://meetingcpp.com/2026/)

**OPEN CALL FOR SPEAKERS**

**OTHER OPEN CALLS**

* **(Last Chance) CppCon Call For Volunteers Now Open** – Interested volunteers have until August 1st to apply at the CppCon main conference which is scheduled to take place from 14th – 18th September. For more information including how to apply visit [https://cppcon.org/cfv2026/](https://cppcon.org/cfv2026/)

**TRAINING COURSES AVAILABLE FOR PURCHASE**

Conferences are offering the following training courses:

**CppCon Online Workshops**

**9th – 11th September**

1. **Modern C++: When Efficiency Matters** – Andreas Fertig – 3 day online workshop available on 9th – 11th September 09.00 – 15.00 MDT – [https://cppcon.org/class-2026-when-efficiency-matters/](https://cppcon.org/class-2026-when-efficiency-matters/)
2. **System Architecture And Design Using Modern C++** – Charley Bay – 3 day online workshop available on 9th – 11th September 09.00 – 15.00 MDT – [https://cppcon.org/class-2026-system-architecture-and-design-using-modern-cpp/](https://cppcon.org/class-2026-system-architecture-and-design-using-modern-cpp/)

**21st – 23rd September**

1. **C++ Fundamentals You Wish You Had Known Earlier** – Mateusz Pusz – 3 day online workshop available on 21st– 23rd September 09.00 – 15.00 MDT – [https://cppcon.org/class-2026-cpp-fundamentals/](https://cppcon.org/class-2026-cpp-fundamentals/)
2. **C++23 in Practice: A Complete Introduction** – Nicolai Josuttis – 3 day online workshop available on 21st– 23rd September 09.00 – 15.00 MDT – [https://cppcon.org/class-2026-cpp23-in-practice/](https://cppcon.org/class-2026-cpp23-in-practice/)
3. **Programming with C++20** – Andreas Fertig – 3 day online workshop available on 21st– 23rd September 09.00 – 15.00 MDT – [https://cppcon.org/class-2026-programming-with-cpp20/](https://cppcon.org/class-2026-programming-with-cpp20/)

**26th – 27th September**

1. **Using C++ for Low-Latency Systems** – Patrice Roy – 2 day online workshop available on 26th– 27th September 09.00 – 17.00 MDT – [https://cppcon.org/class-2026-low-latency/](https://cppcon.org/class-2026-low-latency/)

This is the latest news from upcoming C++ Conferences. You can review all of the news at [https://programmingarchive.com/upcoming-conference-news/](https://programmingarchive.com/upcoming-conference-news/)

**CppCon Onsite Workshops**

All onsite workshops will take place in the Gaylord Rockies in Aurora, Colorado

**12th & 13th September**

1. **Advanced and Modern C++ Programming: The Tricky Parts** – Nicolai Josuttis – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-tricky-parts/](https://cppcon.org/class-2026-tricky-parts/)
2. **C++ Best Practices** – Jason Turner – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-best-practices/](https://cppcon.org/class-2026-best-practices/)
3. **How Hardware Gets Hacked: Breaking and Defending Embedded Systems** – Nathan Jones – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-hardware-hack/](https://cppcon.org/class-2026-hardware-hack/)
4. **Mastering \`std::execution\`: A Hands-On Workshop** – Mateusz Pusz – 2 day in-person
workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-execution/](https://cppcon.org/class-2026-execution/)
5. **Performance and Efficiency in C++ for Experts, Future Experts, and Everyone Else** – Fedor Pikus – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-performance-and-efficiency/](https://cppcon.org/class-2026-performance-and-efficiency/)
6. **Talking Tech** – Sherry Sontag – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-talking-tech/](https://cppcon.org/class-2026-talking-tech/)

 **13th September**

1. **AI++ 101 : Build a C++ Coding Agent from Scratch** – Jody Hagins – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-AI101/](https://cppcon.org/class-2026-AI101/)
2. **Essential GDB and Linux System Tools** – Mike Shah – 1 day in-person workshop available on 13th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-essential-gdb/](https://cppcon.org/class-2026-essential-gdb/)

**19th & 20th September**

1. **AI++ 201: Building High Quality C++ Infrastructure with AI** – Jody Hagins – 2 day in-person workshop available on 19th & 20th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-ai201/](https://cppcon.org/class-2026-ai201/)
2. **Function and Class Design with C++2x** – Jeff Garland – 2 day in-person workshop available on 19th & 20th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-function-class-design/](https://cppcon.org/class-2026-function-class-design/)
3. **High-performance Concurrency in C++** – Fedor Pikus – 2 day in-person workshop available on 19th & 20th September – 09:00 – 17:00 – [https://cppcon.org/class-2026-high-perf-concurrency/](https://cppcon.org/class-2026-high-perf-concurrency/)

**OTHER NEWS**

* **Dates for ACCU on Sea 2027 Announced –** ACCU on Sea 2027 will take place in Folkestone from June 30th – July 3rd with pre-conference workshops taking place from June 28th – 29th
* **Boost Documentary screening at CppCon 2026** – Boost Libraries have announced that they will be screening a documentary on the history of Boost at CppCon 2026. Watch the trailer here [https://www.youtube.com/watch?v=87jvuDbnwqQ](https://www.youtube.com/watch?v=87jvuDbnwqQ)
* **C++Now 2026 Videos Now Being Released on YouTube** – Subscribe to the C++Now YouTube channel to stay up to date when each video is published – [https://www.youtube.com/@CppNow](https://www.youtube.com/@CppNow)

https://redd.it/1v93gvk
@r_cpp
amask = f(amask) on NEON. faster than the obvious blend

## Problem

Apply an operation to elements that satisfy a condition:

for (sizet i = 0; i < n; ++i)
if (mask(a[i])) a[i] = f(a[i]);

### Notes

- `a[i] ∈ (0, 1)`, `thd ∈ (0, 1)`, `mask = a[i] < thd`; uniform distribution (except at the end of the article)
- `f` is one of `sqrt`, `frfrexp` (mantissa), `sin` 3.5 ULP, `sin` 1 ULP, `pow` 1 ULP (from SLEEF)
- `f` and `mask` are passed as runtime values, so they are wrapped in a lambda with `always
inline, otherwise they may
not be inlined
- The array size
n is a multiple of every unroll, tile etc. The tail is trivial to handle(BSL/scalar)
- In tables
` = best, units = GiB/s
- Don't compare numbers across tables. Different conditions, values fluctuate
- All benchmarks: Apple M5; clang++ -O3 -std=c++23 -march=native; GiB/s = (n
4 bytes) / time, min of 720 runs
(During bench, functions run in a changing order, data is restored ofc); n=1e7 + 2432;

## BSL blend

If the problem is memory bound (cheap function or high density), the standard algorithm is optimal:

template <bool Skip>
void bsl(float dst, const size_t n, auto f, auto mask) {
for (size_t i = 0; i < n; i += 16) {
std::array<float32x4_t, 4> v;
for (size_t j = 0; j < 4; ++j) v[j] = vld1q_f32(dst + i + 4
j);
std::array<uint32x4t, 4> m;
for (size
t j = 0; j < 4; ++j) mj = mask(vj);
if constexpr (Skip) {
if (vmaxvqu32(vaddqu32(vaddqu32(m[0], m[1]), vaddqu32(m2, m3))) == 0) continue;
}
for (sizet j = 0; j < 4; ++j) vst1qf32(dst + i + 4 j, vbslq_f32(m[j], f(v[j]), v[j]));
}
}

It computes `f` on every element, but stores only the selected ones.
Skip helps on sparse masks, but otherwise mispredictions will kill performance. We'll need it later.

But for expensive `f` this algo does too much extra work

## Detour

To avoid unnecessary work, we compress selected elements, apply only to them, and expand back.

avx512 does this in two instructions. NEON doesn't, so we'll emulate and optimize.

constexpr size_t tile = 4096;
constexpr std::array<uint32_t, 4> weights{1 + 16, 2 + 16, 4 + 16, 8 + 16};
constexpr auto cps_tbl = compress_table();
constexpr auto exp_tbl = expand_table();
std::array<float, tile + 16> tmp;
std::array<uint8_t, tile / 4 + 3> s;
std::array<uint16_t, tile / 4 + 3> idx; // idx, D and B come in later
constexpr double D = 0.845;
constexpr double B = 0.3;

template <bool Skip>
size_t detour(float
dst, const sizet n, const auto w, auto f, auto mask) {
float* ptr =
tmp.data();
for (size
t i = 0; i < n; i += 16) {
std::array<float32x4t, 4> v;
for (size
t j = 0; j < 4; ++j) vj = vld1qf32(dst + i + 4 * j);
std::array<uint32x4
t, 4> m;
for (sizet j = 0; j < 4; ++j) m[j] = mask(v[j]);

if constexpr (Skip)
if (vmaxvq
u32(vaddqu32(vaddqu32(m0, m1), vaddqu32(m[2], m[3]))) == 0) {
s[i / 4] = s[i / 4 + 1] = s[i / 4 + 2] = s[i / 4 + 3] = 0;
continue;
}

std::array<uint32
t, 4> sk;
for (sizet j = 0; j < 4; ++j) {
sk[j] = vaddvq
u32(vandqu32(m[j], w));
s[i / 4 + j] = sk[j];
}
std::array<size
t, 4> off; off0 = 0;
for (sizet j = 1; j < 4; ++j) off[j] = off[j - 1] + (sk[j - 1] >> 4);
std::array<uint8x16
t, 4> index;
for (sizet j = 0; j < 4; ++j) index[j] = vld1qu8(cpstbl[sk[j] & 15].data());

for (size
t j = 0; j < 4; ++j) vst1qf32(ptr + off[j], vreinterpretqf32u8(vqtbl1qu8(vreinterpretqu8f32(vj), indexj)));
ptr += off3 + (sk3 >> 4);
}
const sizet size = ptr - tmp.data();

if (size == 0) return size;

ptr =
tmp.data();
for (size
t i = 0; i < size; i += 16) {