FITradeoff: the same method, with today's stack
Executive summary
- Why it exists: the method's official system works, but it is old. Built in Delphi, it sometimes freezes or goes offline, and it doesn't use HTTPS. I wanted to show that today's stack can do the same in a way that is more robust, more reliable and better looking.
- What it is: an independent implementation of FITradeoff, which supports decisions with multiple criteria without asking for weights, using only trade-off questions.
- What it isn't: a replacement for the official system. Without access to the original engine, I rebuilt its behavior through reverse engineering.
- Stack: Python with the HiGHS solver, FastAPI, Next.js and an MCP server, all in Docker.
- Quality: about 1,675 tests, with the same result and question sequence as the official system in 67 of 73 comparable cases.
- Time and cost: about 60 hours, half of them on reverse engineering, and about US$1,600 at API prices. That is not what I spent, since I used a subscription plan.
- Status: FITradeoff is now offline, and the code is on GitHub.
Did the summary catch your interest? Then keep going, the rest is for you. If it didn't, feel free to stop here: the essentials are up there, and your time is worth more than my text.
What it cost, in hours and tokens
All of FITradeoff was built with Claude Code, which logs the time of every message and the tokens spent on every response. From those logs, I measured the project's time and cost.
The dollar figure in this section is not what I spent. I used a subscription plan, which is much cheaper. The number shows what the same work would cost paying for the Anthropic API by usage, and it is the closest thing to the real cost of building something like this with AI.
Time
It took about 60 hours of work over 11 days, with 256 requests. Of those hours, between 9 and 19 were mine, reading responses and writing requests; the rest of the time, the agent worked on its own. I only counted stretches with activity, leaving out breaks longer than 10 minutes. Allowing breaks of up to 30 minutes, the total reaches about 82 hours.
Reverse engineering took half of that time, about 34 hours, against 8 for building features. If you are thinking about doing something similar, this is what I would factor into the plan. With a coding agent, implementing what the papers describe goes fast, and what takes time is making the program behave like a system you can only see from the outside. Here, it came to about four hours of comparison for every hour of building, on a system with five modules and 74 compared cases. It is not a rule, and on a larger or less documented system the comparison will likely weigh even more.
Where the time and money went
By type of work, the requests broke down like this.
| What I was asking for | Requests | Time | Cost (API) |
|---|---|---|---|
| Comparison with the official system | 94 | ~34 h | US$641 |
| Refinements and fixes, including code reviews | 55 | ~14 h | US$448 |
| Building features | 41 | ~8 h | US$194 |
| Deployment, infrastructure and security | 46 | ~7 h | US$187 |
| Documentation | 10 | ~2 h | US$59 |
| Environment and tools | 6 | ~1 h | US$33 |
| No clear topic | 4 | ~2 h | US$36 |
| Total | 256 | ~68 h | ≈ US$1,600 |
Comparison with the official system took half of the time and about 40% of the cost, almost as much as building and refinements combined. The section "The papers weren't enough", further down, explains why.
Tokens
It came to about 3 billion tokens, 98.7% of them cache reads, because the model rereads the whole conversation on every response. That rereading costs a fraction of the normal price and, even so, made up about 74% of the bill. The text the model actually wrote, code included, added up to only 6.5 million tokens. Almost everything ran on Claude Opus 5 and Opus 5.5.
How I got these numbers
The prices come from the table Claude Code itself uses in the /cost command, the same as the API's (Opus 5, for example, costs US$5 per million input tokens and US$25 per million output tokens). On one complete session, my figure came out 4% below Claude Code's, so allow a margin of about 5%. The requests were classified by hand, and conversations on claude.ai and requests from other projects were left out.
Why rebuild something that already exists
FITradeoff already existed. The group that created the method, at CDSID/UFPE in Brazil, maintains an official system used in research, and the credit for the method belongs entirely to its authors. What I wanted to put to the test was the stack.
The official system belongs to another generation of software. Built in Delphi, with the IntraWeb framework, it sometimes freezes in the middle of a session, and during my tests it went offline several times, for hours. It is served without HTTPS, so logins and decision data travel unencrypted. The screens exist only in English, with a warning not to use the browser's automatic translation, and criteria show up as C1, C2 and C3. None of this is a flaw in the method; it reflects the time when the software was built.
So the question was whether the same method, built with today's tools, would become a more robust, more reliable and better-looking program. For the answer to mean anything, I couldn't change the method, so nothing counted as done until it passed the comparison with the official system.
The problem with weights
In engineering, it is common to choose between alternatives evaluated on several criteria, and the most widely used tool is the weighted sum. Its problem lies in the weights. It is hard to say with confidence that cost is worth 0.35 and lead time 0.20, and that number often comes from a guess that the whole conclusion then depends on.
FITradeoff, proposed by Adiel de Almeida and colleagues in 2016, replaces the question about weights with comparisons between two concrete situations, in real units, and asks only the questions it needs.
How the method works, without formulas
The person starts by ranking the criteria from most to least important, which already narrows the space of possible weights. Then come the trade-off questions, where the system shows two hypothetical consequences and asks which one they prefer. Behind the scenes, each answer is a step in a binary search over the ratio between two weights.
After each answer, the system uses linear programming to check which alternatives can still be the best under some set of weights consistent with everything said so far. When only one is left, or when the remaining ones are practically equivalent, the questions stop. The implementation covers the five problem types of the official system (choice, ranking, sorting, portfolio and benefit-to-cost). Choice is exact, and portfolio is approximate.
The papers weren't enough
I started with the 2016 paper, the user guide and other work by the group. They explain the method well, but not the rules that make the program behave the way it does, such as the question that opens the session, the order of the criteria pairs and when to stop. Without those rules, an implementation faithful to the papers can reach the same result by a different path, and then it is no longer the same program.
The way out was reverse engineering. I wrote a bot with Playwright that runs the same cases on the official system and on the new version and compares them question by question, and that is how the hidden rules surfaced. On a discrete criterion, such as a score from 1 to 5, the official system only asks at one third, halfway and two thirds between two levels, while my version used a continuous search. It took five rounds for that rule to fall into place.
The bot was a project of its own, because the official system rejected automated browsers, had buttons that disappeared mid-click and sessions that froze for hours. That is why every measurement was recorded, and the parity suite runs offline. Some points, such as the formula for the veto model, could not be confirmed and are recorded as inferences.
The two versions, side by side
The images below show the same steps in both systems, with the official one on top. The screenshots of the official system come from FITradeoff.org, by CDSID/UFPE, and are here for comparison only.
The home screen
Both screens offer the same options, including the two portfolio variants. In the official system, the modules appear with their English names and a Portuguese translation in brackets, and only the portfolio has a help button. In the new version, each card says how that decision ends, the next steps appear at the top, and the data source, by hand or by spreadsheet, is chosen right there.


Performance matrix


Trade-off question


Hasse diagram
When the order is not settled yet, both systems draw which alternative never falls behind which. In the compared case, the dominance pairs come out the same in both.


The stack
The repository has five parts: the engine (engine/), the API (api/), the MCP server (mcp_server/), the interface (web/) and the parity suite (paridade/). The engine uses Python 3.13, NumPy, SciPy and the HiGHS solver, the API uses FastAPI, and the interface uses Next.js 15, React 19, TypeScript and Tailwind. Everything runs in hardened Docker containers that restart on their own, behind Caddy with HTTPS, and deployment through GitHub Actions rolls back to the previous version if it fails. A test also reads the engine's code and fails if one layer imports another it shouldn't, so the architecture doesn't depend on discipline alone.
Storing only what the person said
The state of a decision is not stored, only the list of what the person declared (the order of the criteria and each answer). Everything else is recomputed by replaying that list. It looks wasteful, but it was one of the choices that paid off the most. Undoing an answer means cutting the end of the list, contradictory answers show up as an infeasible linear program, and switching problem types doesn't require asking everything again.
Asking fewer questions was not the goal
I tested an information-gain question selector that, in a ranking case, solved in 10 questions what the official system solves in 21. The number looks great, but that selector pulled five cases that already matched away from the official system, so I kept the one that reproduces the method's path. The thresholds that decide when a pair of criteria is exhausted also came from measurements against the official system, not from guesswork.
The assistant decides nothing
With the MCP server, you can run a decision by talking to Claude, but the server only translates the protocol and computes nothing on its own. The assistant is instructed never to answer on the person's behalf, and the answer tool requires the question's identifier, so no answer gets recorded for a question nobody saw.
Three measurements that changed the code
The solver. The API calls HiGHS directly, with one solver per request and without reusing the basis between queries, which used to make the same decision give different results. With sixteen simultaneous connections, that gave 56.8 requests per second, against 16.3 with SciPy.
The queue. Each heavy computation waits for a slot before taking a thread, with at most two per person and 24 in total. Without it, one person with several large decisions could freeze everyone else, and another person's health check went from 5 milliseconds to 68 seconds.
Memory. Linear algebra libraries open one thread per core. Pinning each of them to a single thread, with three environment variables, cut the process memory from 1.5 GB to 124 MB.
Security
FITradeoff lived behind FasorX, which handled login and issued a 5-minute token signed with RS256. Without the issuer and audience configured, the process wouldn't even start, so the dangerous state, published and unauthenticated, couldn't be reached by forgetting a setting. Each person's directory is a SHA-256 digest of their identity, and the part exposed to the internet is a separate application, with only /mcp and the health check.
Assistant keys have 256 random bits, expire in 90 days and only their digest is stored. Since they are random, a slow KDF like bcrypt would only add cost. A security scanner, OWASP ZAP, also led the interface to build its own CSP, with a fresh nonce on every request.
How to know it's right
Besides parity with the official system, there is an oracle test, in which a simulated decision maker with hidden weights answers the questions, and the test fails if the method discards the alternative that is actually best. Even documentation is enforced by a test, which fails if a relevant function doesn't state the why, its source in the method and whether the result is exact or approximate.
Where the new version falls short
It would be dishonest to finish without this part. The new version does not replace the official system. Everything it knows about the official behavior came from observation, which covers the cases tested, not every possible one. With access to the original engine, the right path would be to keep it and replace only what surrounds it.
Parity is not complete either. Of the 6 cases that don't match, 4 happen when the first question is answered with B, a situation in which the official system sometimes repeats the previous pair. I found a pattern that reproduces those four cases, but I didn't implement it, because it is not a rule I understand. The other 2 end one question apart, on a pair that sits less than 0.00005 from the stopping threshold.
The official combinatorial portfolio could not be compared, because its result never became available in my tests, and the new version solves it approximately. Finally, the new version was online for a short time, so I can't claim it would be more available over the years, only show how it was built for that.
If you want to run it
The code is in the repository. With Python 3.13 and Node.js installed:
python -m venv engine/.venv
engine/.venv/Scripts/python -m pip install -e engine[dev] -e mcp_server -e api
cd web && npm install
On Windows, .\iniciar.ps1 starts the API on 127.0.0.1:8000 and the interface on localhost:3000. Tests run with pytest engine/tests mcp_server/tests, and the parity suite, offline, with python -m pytest paridade/. To go deeper, the repository has the lessons from reverse engineering, parity with the official system and the measured behavior, case by case, all in Portuguese.
To wrap up
FITradeoff did what I wanted from it. Rebuilt with today's stack, the same method came to run with HTTPS, with tests that compare it to the original and with clearer screens, reaching the same conclusions in 67 of 73 comparable cases. That is why it could go offline, and the code is still available. If you use FITradeoff and want to talk about it, reach me through the contact page.
This FITradeoff is an independent implementation, with no ties to the group that created the method, and its values may differ from the official system. It was built from the practical guide, the published papers and observation of the official system.