PDF, spreadsheet, Word, CSV and SQL dumps enter the project and are read for real: a profile over the whole file, a role declared by whoever uploads, and an explicit warning when a file came in empty, truncated or looking like an error.
The problem
The project’s context is scattered across attachments no tool reads in full. The spreadsheet enters as chopped text, the vendor’s proposal is treated as truth, and a corrupted file is politely summarized as if it were data.
Each document enters with a role declared by whoever uploads it: snapshot of the environment, third-party proposal, requirement. A large CSV comes back as a per-column distribution over the entire file, and a degraded response is flagged to the consultant, never disguised.
How it works
The same PDF page can be your real environment or a vendor’s proposal. The product does not guess: the uploader declares the role, and a third-party proposal enters marked as not endorsed, to be challenged rather than absorbed.
The declared context travels along: "the as-is up to page 11 is real; from there on it is a proposal". That is exactly the kind of boundary that separates analysis from guesswork.
Documents
Samples, schemas and specs that EVODA reads and uses across every engagement.
| File | State | Declared role | EVODA’s summary |
|---|---|---|---|
| schema_producao.sql | ready | Snapshot of the current environment | 48 tables, 9 without a primary key, 3 personal columns unclassified. |
| proposta_lakehouse.pdf | ready | Third-party proposal, not endorsed | Proposes a full platform; volume assumption 20x above what was measured. |
| vendas_2025.xlsx | processing | Reference material | generating… |
Whoever uploads declares what the file is. The role is never inferred from the content or the name.
A spreadsheet becomes text per sheet, with formulas delivering their results; a large CSV comes back as per-column distribution, cardinality and extremes computed over the whole file, with truncation declared when it exists.
A file that is empty, shrunken, duplicated or looks like error output is flagged as such. The consultant decides what to do with that; what it does not do is summarize garbage as if it were context.
Now: alternatives compared against weighted criteria, cost computed per tool, assumptions on record.
We want to move everything to a data warehouse. Which one do you recommend?
Before choosing, I need two numbers I do not have yet: your active volume and how many people query per day. You mentioned 40 GB and 25 BI users. At that size, a warehouse usually costs more than it solves.
I will record that as an assumption, with the condition that would invalidate it: if active volume passes 1 TB, the math changes.
Makes sense. Compare the two options then.
Requirements captured
Assumption on record: "active volume stays below 1 TB". Re-checked on every read of the environment.
Under the hood
The ingestion rules that apply to every document in the project.
What comes next
Discovery happens through guided investigations: read-only scripts you execute, with the result coming back tied to the question.
See it from the inside →ConsultantGuided interviewDiscovery is led by the consultant: one question at a time, disagreement backed by evidence, and every answer recorded as a requirement.
See it from the inside →LGPDLGPD per columnEvery column with a classification and an anonymization technique, views for non-production environments, and the impact report when the design demands it.
See it from the inside →Paste the structure, get the most serious findings in about two minutes, and decide later whether an engagement is worth opening. No account, no card.