I use conditional formatting to color cells according to the probability that I can lift them—if I lifted 50kg for 10 reps then I can definitely do 50kg for 9 reps, so that cell is green. But if e1RM(50,10) > e1RM(40,15) then I can probably do that too so it's light green. The visualization naturally becomes Pareto-like.
If I'm feeling strong I can aim for higher weight, lower reps. Or if I'm feeling weak I can close out a (weight, reps) that's below my current e1RM but I haven't accomplished yet. The end result is that I'm always "accomplishing" some sort of PR no matter how I feel.
I call this e1RM Bingo.
When explaining it to some coworkers, I stumbled on a fairly intuitive explanation: "I've run farther before, and I've run faster before, but I've never run _this_ far, _this fast."
There was some pushback about why not just call it a PR (personal record), but I would only use that term for fixed distances (1mi, 5k, 10k, etc.) or a consistent route that I've run many times before. Nobody would say "I set my 7.40 mile PR today." More importantly, it misses the comparison to all farther (and faster) runs—it's not exciting to set a 5k PR just because you've barely run that distance before, and the pace is actually slower that a 10k you've done.
(Had a Pareto run of 7.40 miles @ 6:28/mi last week!)
ChatGPT 5.6 Luna on the right (cheaper) cover most of the frontier, with a point for Deepseek flash, and higher performance overlapping heavily between 5.6 Sol and Fable.
That DeepSeek point will probably move back towards Luna as deepseek announced a "significant" price increase coming to their API [1], which kind of demonstrates that beating the Pareto frontier is where the difficulty actually is).
[1] https://www.bloomberg.com/news/articles/2026-08-06/deepseek-...
Maybe in vein but did anyone already figure this one out? The closest I got was PT sans, open-licensed commissioned by the Russian ministry for communication (I found it surprising that a country that doesn't use Latin script made the best font!), but it's not widely shipped so you need to figure out how to include font files whenever you want to use it
My main finding for “pick whatever weight you want today” was that picking a lot of different weights made the curve less identifiable, so my latest iteration encourages you to pick a ladder for a few sentinel exercises per mesocycle in order to improve the statistical power. In addition, strength improves more quickly at >80% of 1RM, and hypertrophy depends on proximity to failure, so if you pick a lower weight, you really need to go to failure, which burns you out for the rest of your session, where leaving 1-2 reps in reserve is probably sufficient for hypertrophy and leaves a lot more gas in the tank for the rest of the session. Definitely open to suggestion/discussion here.
https://curvefit.app (it runs on Cloudflare free tier, so I won’t have to start running ads or charging until I hit a couple thousand users)
Please don't make an app based on this.
If one option is at least as good on every relevant dimension and better on one, just pick it. That's not really a trade-off, and it shouldn't need escalation. Eg, if two SaaS tools cost the same and have similar support, but one fits your use case better, you choose that one. Otherwise, you just suck at your job!
The interesting decisions only start once you're already on the frontier, where getting more of one thing means giving up something else. If the better tool costs 50% more, now you're trading capability against cost, and that may need sign-off.
Basically, everyone should be able to get to the frontier on their own. Coordination and arbitration at higher levels of the org / between different departments should happen on the frontier, where the trade-offs involve several people or teams.
Example: Which LLM gives me the best ELI5 explanations for a given price. https://evalry.com/benchmarks/explain-like-i-m-5-321
Matthias Ehrgott's books on multicriteria optimization explain Pareto efficiency very well without sacrificing rigor. I think they do a better job than this article.
The 80/20 “rule,” as far as I know, is meant to be descriptive after the fact. It can’t be used as a planning assumption. To be fair to those managers, they don’t really mean to be rigorous. They are just trying to justify cutting scope.
As the number of objectives (dimensions) increases, the number of samples you need to cover the frontier increases exponentially. You will very rarely find solutions that actually dominate other solutions in many practical optimization scenarios. With 2 dimensions you have a 25% chance of domination. With 10 dimensions it's a .098% chance.
The most useful cases I've seen tend to occur where we just optimize for two things at once. The chances of domination are high, it's easy to visualize and very efficient to implement. As we get into higher dimensional spaces, things get weird really fast.
But they go even a step further, they extend into 3 dimensions to also add body weight as a variable. So your graph would really have to be a 3D volume. Because different levels of body weight have different capabilities.
I think it's great and hope the price can stay the same.
If you want the most 'optimal' form of this (aka, hell on earth), you should purchase a rowing machine. Being able to engage with very aggressive, full-body exercise every single day without exceptions is almost like cheating biology. You can maintain a 2-3x VO2 max premium over your peers with very little risk of injury.
Some nuance here: the latest research shows that proximity to failure is the main hypertrophy driver regardless of load and rep count; high rep count makes proximity to failure harder to gauge; so high load/low reps close to failure is probably better for hypertrophy (there are other good reasons to do higher reps/lower load work though)
If pursued, good luck!
Anyways I'll namedrop Iosevka as perfect monospace font for working on 13" laptop
"The Pareto Front today claimed responsiblity for...."
The geometric problem of computing a d-dimensional Pareto set of cardinality n
https://en.wikipedia.org/wiki/Maxima_of_a_point_set
has a truly weird property not covered by the computational complexity discussion on that page. It says there's an algorithm achieving O(n log(n)^(d-3) log log n), which is true and also a lie. The algorithm that achieves that asymptotic form is a galactic algorithm; and not an ordinary one in the sense of "has a large constant multiplicative factor", but one with this property (I've never found any other algorithm which exhibits it):
The runtime is within a bounded constant factor of n^2, for all n up to some critical N whose size is exponential in d (I think it was exactly 2^d or something).
I.e. the runtime has "two shapes": it's purely quadratic up to a galactically-large constant, and thereafter has a transition into to a slower function. The asymptotic version in the textbooks isn't achievable in the real world (for all but very small dimension).
There's an elementary proof using generating functions.
edit to add: If anyone's curious about it, a simplified version of the recurrence relation that's enough to exhibit this behavior (you can instantly see it if you graph this numerically) is
f(n,d=0) = 1
f(n=1,d) = 1
f(n,d) = n + 2f(⌊n/2⌋, d) + 2f(⌊n/2⌋, d-1)Is your planet accepting immigrants? I think I'd like it there
I've built large, deep product evaluation frameworks, and it is 100% of the time a running argument with stakeholders, inside and out, "well you should have measured it this way" or "I think we should be targeting X not Y" or "why didn't you consider Z in the metric??"
The Pareto Front in practice is squishy, fuzzy, and often quite moist and moldy.
As you say, the most useful things happen in low-dimensional spaces.
Some thing is "pareto optimal" when there isn't another thing that's AT LEAST AS GOOD in ALL measures, and BETTER in at least one way. For example, if we say there are no ties (for simplicity), then the cheapest language model is pareto optimal; the fastest model is pareto optimal; those which score highest on each benchmark are pareto optimal; and so on.
Tradeoffs can also be pareto optimal: for example, if the cheapest model is also slow, then there will be more pareto optimal models which are "cheapest for their speed"; and so on for other tradeoffs (e.g. fastest that achieves a certain benchmark score; cheapest model with open weights; etc.).
If you're making a decision about which thing to choose, you only need to care about those in the pareto front (since, by definition, anything that's not pareto optimal is objectively worse on at least one measure).
Pareto optimality does not compare one measure against another: something that's 10000x slower can still be pareto optimal, if it's 1% cheaper than the alternatives. To pick a "best" thing, you could give a weight/importance to each measure, and combine them into an overall score: but that's subjective, and might vary between people and tasks. In contrast, focusing on the pareto front is a way to ignore those things that will never be the best, regardless of weighting.
“Chapter 4: The Seam Model”, Michael C. Feathers, Working Effectively with Legacy Code
Of course, what's hard anyways when you have a good set of solutions that are pareto optimal, is to then choose between them. Especially as the dimensions (objectives) grow. In my example we can end up with many variants of strength/weight trade-offs that each are optimal, which one to choose?
Say a race vehicle has acceleration, top speed as defining parameters. Some are slow but accelerate hard, others need a long time to reach very high top speeds. Others are in between, or just flat out bad at both.
The pareto frontier is the set of vehicles that are best: pick one from the frontier and you can be sure that for it's given top speed, none accelerate faster. And vice versa, pick one with a given acceletation and you are sure none have a better top speed
They apparently released Hyperlegible Next in 2025 which, flipping between tabs on Google Fonts (since the original website doesn't show the fonts), is nearly identical but has five new weight settings (nobody should imo ever use thin fonts though, it noticeably harms readability for me and my sight is only the tiniest bit below normal vision, but ok it's an option) and improved kerning (the original font had extremely little space between 'll', for example)
The 2025 version sadly doesn't ship with my version of TexLive, but the original (from 2020) already does so that makes it easy to use as well! Cool stuff, thanks for the tip :)
ss02 disambiguation seems to be the one I'd be wanting to turn on, with tnum for monospace numbers being a good option as well that I hadn't even realised I wanted from a font!
It’s really that simple.
Eschew obfuscation.
Going from 81GB of weights to 79GB of weights can mean a 50% reduction in GPU capacity required.
If you can fit a model in just one GPU (or rack) as opposed to across an entire datacenter, the latency gains can be substantial too. If you can reduce token latency by half, that would double the amount of customers you could support.
Mapping the cost of something (like an algorithm), and the time it takes (so lower is better for both). 1, 3 and 5 are all optimal in their own sense. No one is strictly better than the other, just different tradeoffs you have to choose yourself. However, you would never choose 2, because for a lower cost you could get the same result choosing 3. Same with 4, 6 and 7, they all have something that's both faster and at the same time just as cheap you could choose.
A pareto front is a bit like the classical "fast, cheap, good, choose 2". There are always tradeoffs, but if something is both slow, expensive and not better than something that's faster and cheaper, it's a bad choice, and thus not on the "pareto front".
[1] - http://montypython.50webs.com/scripts/Life_of_Brian/8.htm
Tabular numbers are awesome!
The Pareto points are where you sacrifice the least of anything to get the most of everything.
There's the saying about buying computers. Good, Cheap, Fast, pick any two. That's where you would prioritise.
If someone makes something that better, cheaper, and faster, or even pretty close to the best on two of those and clearly better on the other. It's a Pareto point.
Over time computers are getting better, cheaper and faster (software notwithstanding). The leading edge of that advance of all of the things is the Pareto front.
> a Pareto front represents the set of solutions where no solution outperforms any other solution in the set at every objective
I do not believe you are correct when you say
> something that better, cheaper, and faster, or even pretty close to the best on two of those and clearly better on the other. It's a Pareto point.
Since that would outperform on every objective
GP's point that it's prioritisation does not seem incorrect to me. Prioritisation involves considering trade-offs of various approaches and deciding which aspects & attributes to optimise for, at the expense of others.
We choose our items/workflows/technologies/whatever, so we get the best/most efficient/most effective/whatever, across the widest possible set.
Sounds like prioritizing, to me, but I’m just a dumb hick, so I suppose I can be wrong.
Going for the Pareto is when you elect not to prioritise. It is explicitly deciding to not choose one property over another ant to keep everything as much as you can.
Getting to that point can be calculated (in some cases), but I suspect most folks get there by trial and error. Finding out what is effective, and what is not, and choosing what is effective, over what is not, until there's no longer a choice. That often becomes tribal knowledge, and is handed down. There's always someone trying to improve it, and when they figure it out, that gets added to the tribal knowledge. Basically, that's how nature does it, so there's some serious prior art. Natural Selection is brutal prioritization.
In Morocco, they used to announce the end of the Ramadan fast, by holding up a black thread and a white thread, and waiting until they could not tell the difference.
Then, they'd fire a cannon, and everybody would dig into some awesome soup. Sort of the same thing.
From Wikipedia, the free encyclopedia
In multi-objective optimization, the Pareto front (also called Pareto frontier or Pareto curve) is the set of all Pareto efficient solutions.[1] Colloquially, this means when there are many distinct objectives to consider in an optimization problem, a Pareto front represents the set of solutions where no solution outperforms any other solution in the set at every objective, and every solution not in the set is outperformed by at least one solution in the Pareto front in every objective.[2] The concept is widely used in engineering.[3]: 111–148 It allows the designer to restrict attention to the set of efficient choices, and to make tradeoffs within this set, rather than considering the full range of every parameter.[4]: 63–65 [5]: 399–412
Example of a Pareto frontier. The boxed points represent feasible choices, and smaller values are preferred to larger ones. Point C is not on the Pareto frontier because it is dominated by both point A and point B. Points A and B are not strictly dominated by any other, and hence lie on the frontier.
A production-possibility frontier. The red line is an example of a Pareto-efficient frontier, where the frontier and the area left and below it are a continuous set of choices. The red points on the frontier are examples of Pareto-optimal choices of production. Points off the frontier, such as N and K, are not Pareto-efficient, since there exist points on the frontier which Pareto-dominate them.
The Pareto frontier, P(Y), may be more formally described as follows. Consider a system with function , where X is a compact set of feasible decisions in the metric space
, and Y is the feasible set of criterion vectors in
, such that
.
We assume that the preferred directions of criteria values are known. A point is preferred to (strictly dominates) another point
, written as
. The Pareto frontier is thus written as:
[edit]
A significant aspect of the Pareto frontier in economics is that, at a Pareto-efficient allocation, the marginal rate of substitution is the same for all consumers.[6] A formal statement can be derived by considering a system with m consumers and n goods, and a utility function of each consumer as where
is the vector of goods, both for all i. The feasibility constraint is
for
. To find the Pareto optimal allocation, we maximize the Lagrangian:
where and
are the vectors of multipliers. Taking the partial derivative of the Lagrangian with respect to each good
for
and
gives the following system of first-order conditions:
where denotes the partial derivative of
with respect to
. Now, fix any
and
. The above first-order condition imply that
Thus, in a Pareto-optimal allocation, the marginal rate of substitution must be the same for all consumers.[7]
Algorithms for computing the Pareto frontier of a finite set of alternatives have been studied in computer science and power engineering.[8] They include:
Since generating the entire Pareto front is often computationally-hard, there are algorithms for computing an approximate Pareto-front. For example, Legriel et al.[19] call a set S an ε-approximation of the Pareto-front P, if the directed Hausdorff distance between S and P is at most ε. They observe that an ε-approximation of any Pareto front P in d dimensions can be found using (1/ε)d queries.
Zitzler, Knowles and Thiele[20] compare several algorithms for Pareto-set approximations on various criteria, such as invariance to scaling, monotonicity, and computational complexity.