I got a peek at our AI operations dashboard last week, care of Neal Iyer (thanks Neal). Like many of us, we are all trying to get a handle on the economics of AI. We already know that Tanya Brennan can spend $1,000 on Fable in one week without thinking about it. And we know that if we're not careful, our Amazon Web Services (AWS) Bedrock bill can double in a month. These cost components are the easiest to at least make visible. The total cost of AI, and the extent to which these costs are creating a return, continues to be a frustrating topic of conversation.
There is no one size fits all model for evaluating the cost of AI solutions and vendor pricing methodologies are changing. What is not changing (at least for the next 42 seconds) are the ten cost dimensions. I've grouped these into four buckets, and each dimension should be looked at for every major AI investment. A subset should be looked at for individual use cases, or small groups of use cases.
Buy v. Build
The most obvious cost elements are, of course, the application itself and the any use charges. We are pretty familiar with these as they are part of the traditional software-as-a-service (SaaS) model. This can be convenient pricing model with a good partner as it creates a predictable cost pattern. Whether you buy or build, you will likely have to integrate, which is your cost to deliver.
The choice of buy versus build has an impact on the downstream cost dimensions like talent. When you build, you have to get more of a very specialized type of talent that is difficult to find. Growing this talent internally is expensive and time consuming, partnering to acquire the talent is an option. Regardless, it is essential that every mortgage organization have some subset of W2 talent for AI and generative AI.
Variable Costs
Oh the tokens. So much controversy, so much variability. A token is a word or part of a word, and it is the currency of large language models (LLMs). You pay for tokens as they go in, "thinking tokens", and the tokens that come out. Even if they are wrong.
The good news is that the cost of foundation AI has decreased significantly since "the early days" of 2023 and 2024. The price of a fixed level of AI capability has fallen 9x to 900x per year depending on the task (Epoch AI, 2025). GPT-4-level output that cost ~$30 per million tokens in 2023 now costs well under $1 (industry pricing analyses, 2026).
The bad news is that the cost of LLM usage for most organizations is still going up. Why? Because we typically use the newer bigger models for everything, without optimizing, and the price for those models has not gone down.
We recently implemented our AI ops dashboard for cost visibility, and have a range of techniques we use for cost optimization, including model distillation, and intelligent design of our innovation pipeline. Model distillation is the process of learning from a higher horsepower model and then training a lower cost model to act like the bigger one. Intelligent design in the innovation pipeline is the process of dynamically optimizing model choices and context management in the pipeline relative to the tasks they need to perform.
Connecting Tokens to Outcomes
The next big challenge for us, and really for everyone, is how to connect token use to outcomes. We get the big bill from AWS, and we can connect that bill to people and to models, but we don't know if the spend was "worth it". This comes down to LLM observability and monitoring, and it's not easy but it can be done.
Step 1: Find the Spend
The first step is to find all the spend - subscriptions, cloud charges, vendor products, API charges - finding and gathering the cost for every person or team using generative AI. This sounds easy, but really isn't. For example, this may mean combing through credit card statements for that pesky Synthesia subscription you forgot you had. And then you have to do this across all credit cards.
Step 2: Connect the Spend
This step requires you to connect the dollars spent to the people who spent them. You found that $2,500 charge for Replit in May - now you have to connect that to the people who have access. You have to find the individuals who are using the product compared to ilde seats and users. This requires you to go into yet another platform, and connect that platform to you AI operations solutions. This can make your team nervous. People generally do not like to be monitored. And the point is not to monitor the people, it is to understand how AI is being used for the purpose of optimizing cost. You may even need some teams or people to spend MORE.
Step 3: Connect the Intent
This is the hardest part, and requires an even lower level of monitoring as you want to see not only THAT the people are using AI but WHAT they are using it to create. If you only use one platform, say Microsoft Copilot, this won't be that hard. But if you use a lot of different platforms, this can be a challenge. We have recently implemented Langfuse for LLM observability and have started to connect prompts to tokens, models, tool use, and agent invocation.
Step 4: Elevate Usage to Outcomes
So now you found all the spend, connected it to your AI ops console, and are able to observe the relationship between the dollars and the purpose. Next, and most difficult, you elevate that to the outcome level. Full transparency, we are still working on this. We have pockets of use that we understand well, and pockets that we still have to connect. Spend + person + intent = the basis for understanding the outcome.
The Cost of Compliance
This is your data, your evaluation process, and your compliance approach. Much has been said about the important of good data, I won't dwell on this here. But I will talk about the other costs of quality - compliance and evaluation. Cost of compliance is the cost of the program and processes you have in place to govern your AI use - and especially your high risk uses.
Step 1: Catalog Use Cases
Step one is to understand what you are doing with AI and does any of it qualify as what Colorado is calling a consequential decision. I use this standard because it's the best one I can find, and it works well with what has been defined by the Administration for compliance by the federal government. The decision to place something in the catalog as well as the decision to designate a use case as consequential or high risk should be documented.
Step 2: Validate Your Governance Program
I am advising everyone to put in place the governance model proposed by the Administration in M-25-21, I think it's a solid framework that is straightforward. The hardest elements are the logging and monitoring, both of which needs to be in place for any AI use case in my opinion.
One thing to note, a human-in-the-loop who just reviews AI generated content will ultimately be insufficient to meet what I think examiners and regulators will look for. In addition, you will want to understand, document, and be able to explain and demonstrate effectiveness of your context specific guardrails.
Step 3: Prioritize Your Evaluations
In addition to guardrails, another key cost of compliance is your evaluation program. These are the tests, preferably performed at scale by automation, that you perform to test that AI systems are performing in accordance with your expectations - including standards for quality. If you do not have an evaluation program, start small, with core metrics that matter.
Step 4: Create Your Benchmark Data Sets
There really is not substitute for high quality benchmark data sets. These are sometimes also called QA (question and answer) pairs. QA pairs are are structured data sets documenting representative test questions along with their verified, correct reference answer (often called "ground truth"). They are used as a standard benchmark to test how accurately and safely an AI model or retrieval system responds. These are time consuming to do well but will be a critical part of your ongoing quality process.
Step 5: Implement Your Evals
You want to run your evals on a schedule, anytime you need to test prompt changes, and then before and after model upgrades. Depending on what information your vendor makes available to you, you might be able to rely on their evaluation program but at the end of the day, you are accountable. As always, I recommend the open source framework by Promptfoo
This is a sample eval dashboard from promptfoo, I find this visual helpful to understand what an eval is and why it's so helpful.
Step 6: Monitor and Remediate Results of Your Eval Program
As you run your evals, you might see that your eval scores are getting worse, you will need to take action if this is the case. Your eval program will be a critical investment in the robustness of your overall AI ecosystem and should not be skimped on. It may end up being a significant cost to consider.
The Cost of Innovation
And finally, the cost of innovation. This article is already way too long so I'll summarize. Talent, time, change, and opportunity cost. These are the hardest to quantify and the most important to do well. I promise to write another article on this specifically as part of this series.
PhoenixTeam Launches Human Futures, First Cohort Begins Hands-On AI and Career Readiness Work
ARLINGTON, VA, UNITED STATES, June 9, 2026 -- PhoenixTeam announced the launch of Human Futures, a new initiative helping young people prepare for a world being reshaped by artificial intelligence. Human Futures is grounded in a single belief: the next generation will inherit a different world than the one we grew up in, and it is our responsibility to prepare and employ them. Through AI education, career readiness, employer-connected learning, and hands-on project work, the initiative helps students and early-career professionals develop the skills, confidence, and judgment to use AI thoughtfully, solve real problems, and shape what comes next.
For PhoenixTeam, Human Futures is a natural extension of the company’s work helping organizations understand and apply AI in practical, responsible ways. After years of building AI-enabled solutions, educating mortgage and technology professionals, and helping clients navigate generative AI adoption, PhoenixTeam is now applying that experience to a broader challenge: helping people prepare for the world AI is creating.
The first Human Futures program is underway through PhoenixTeam’s 2026 Summer Internship Program. The cohort includes interns working across Human Futures, AI services and Phoenix Burst, project delivery, business operations, and marketing. Throughout the program, interns receive weekly AI education, learn how businesses operate, build communication and career-readiness skills, and work on individual and team projects that will become part of a real portfolio.
“All of the content we deliver in our Human Futures programming will be accessible to all and available for free. We hope this will be the first of many ways that we can use fear to catalyze action,” said Tela Gallagher Mathias, CTO of PhoenixTeam and CEO of Phoenix Burst. “Human Futures is about bringing the focus back to people. We want young people to understand AI, work with it responsibly, and still build the human skills that make them valuable: courage, judgment, aspiration, resilience, and authentic human connection.”
“The future of work is changing quickly, but preparation cannot just be about tools,” said Tanya Brennan, CEO of PhoenixTeam. “Human Futures is about helping people build the thinking, confidence, and real-world experience they need to participate in that future. The first cohort is already showing what happens when young people are given real problems, real mentors, and room to build.”
Human Futures will continue to grow through programs for students, early-career professionals, educators, families, and employers. The initiative will focus on practical AI readiness, responsible use, portfolio-based learning, and employer-connected pathways that help participants move from learning to contribution.
About Human Futures Human Futures is a PhoenixTeam initiative preparing learners to think, work, and lead in a world being reshaped by artificial intelligence. Built for learners, parents, educators, and employers, Human Futures connects AI education with real-world projects, mentorship, career readiness, and portfolio-based learning. The initiative is rooted in a human-first belief: AI should expand what people are capable of, not replace the thinking, creativity, judgment, and resilience that make people valuable. Through hands-on programs and employer-connected pathways, Human Futures helps learners build practical skills, confidence, and real experience for the future of work and life. For more information, please visit www.aifutures.com
About PhoenixTeam PhoenixTeam is a woman-owned technology services firm headquartered in Arlington, Virginia, specializing in AI-powered mortgage operations and technology services for the mortgage and financial services industries and federal housing agencies. Our mission is to enable affordable and accessible homeownership through innovative, customer-centric technology. With a strong focus on generative AI, we tackle complex industry challenges, equipping businesses with cutting-edge tools that enhance innovation, efficiency, and compliance. By bridging the gap between technology and business teams, we strive to bring joy and purpose back to software development, making a meaningful impact in the lives of our clients and homeowners everywhere. For more information, please visit www.phoenixoutcomes.com.
Phoenix Burst Wins MortgagePoint Tech Excellence Award for Second Consecutive Year
Phoenix Burst, PhoenixTeam’s AI-powered regulatory intelligence platform, has been recognized as a winner of the MortgagePoint Tech Excellence Award, honoring the most innovative technology providers transforming the mortgage and real estate industries. This marks the second consecutive year Phoenix Burst has received this recognition.
Phoenix Burst was built to simplify how mortgage organizations manage regulatory change. The platform identifies federal and state regulatory updates, creates clear change statements, and produces delivery-ready requirements, user stories, acceptance criteria, and test cases. With guided human review and traceable outputs, Phoenix Burst helps legal, compliance, product, operations, and delivery teams reduce manual effort, standardize interpretation, and move from analysis to implementation faster.
The platform’s design prioritizes both innovation and responsibility. Built-in safeguards including human-in-the-loop curation and retrieval-augmented generation ensure that AI-driven efficiency aligns with compliance requirements and industry standards. Phoenix Burst is also SOC 2 Type II compliant. That standard means an independent auditor tested the platform's data security controls over time and confirmed they operate as intended. For enterprise mortgage organizations, that verification confirms the platform protects sensitive information with the controls their compliance teams require.
This recognition affirms a belief at the center of PhoenixTeam's work. Innovation and responsibility can advance together. By removing the manual effort that once slowed regulatory change, Phoenix Burst returns time to the teams that need it and positions the mortgage industry to lead rather than react.
PhoenixTeam Earns Inc. Best Workplaces Recognition for People-First Culture
PhoenixTeam has been named to Inc. Magazine’s 2026 Best Workplaces list, the publication’s annual recognition of companies building cultures that keep and inspire their people. The honor reflects what PhoenixTeam’s own employees say about a workplace where they feel valued, supported, and able to do their best work.
Inc.’s Best Workplaces recognition is earned, not claimed. Companies are evaluated through anonymous employee surveys that measure key aspects of the workplace experience, including engagement, management effectiveness, professional development, work-life balance, and benefits. That commitment to employee well-being is reflected in offerings like fully employer-paid family health care, a benefit that changes the lives of employees and the families who depend on them.
The recognition reflects a culture intentionally built around employee growth, well-being, and connection. Guided by six core values, lifetime learning, impact, entrepreneurialism, service, diversity, and wellness, PhoenixTeam invests in opportunities that help employees thrive.
This award affirms what PhoenixTeam employees experience every day: a workplace worth choosing and a company committed to investing in the people who make its work possible.