Charlotte, NC
ProjectsOctober 9, 2026

Solomon Thrift Store Pricing

Tap image to enlarge
A local Christian ministry runs a thrift store, and almost everything on its shelves arrives as a donation. Someone has to decide what each item is worth, and that pricing job falls to staff and volunteers who are not resellers. They do it on a phone. The store wants a fair price that sells within a couple of weeks. It doesn't want the highest price eBay might bear, and it doesn't want to give things away. Every tag is money for the ministry. The workflow I set out to replace was simple, and it worked. Someone took a photo, pasted it into a general chatbot, and asked what to charge. It was fast, and people found it useful. It also left nothing behind. Nobody could see the basis for an answer, a manager had no chance to catch an expensive or risky item before it hit the floor, and there was no way to tell whether prices were drifting. Solomon is the thrift store pricing app I built to replace that habit. A volunteer photographs an item and adds a short note. Solomon researches it and recommends a shelf price with an explanation, a confidence level, and a checklist of things to verify in person. The volunteer accepts the price, adjusts it, or asks a follow-up question. Expensive, uncertain, and sensitive items go to a manager first. Admins tune those rules, and everything lands in a searchable history. It's named after the king known for good judgment, and the interface never says "AI." The app researches an item. That follows the ministry's preference, and it keeps price notes reading like a person wrote them instead of like a research transcript. The server strips URLs, markdown, and citations from every free-text field the model writes. The constraints shaped every decision. Volume is low, the budget is close to zero, and the people using it are on phones and are not technical. Solomon has been in use at the store since the end of July. By the start of October, staff had made final pricing decisions on 184 items through it. The more interesting story is what happened after launch. I spent two months making the pricing pipeline more rigorous, one real failure at a time. Then a backtest against those 184 staff decisions showed that a much simpler design priced a little more accurately and without production's upward bias. Most of this page is about how I got there. Solomon is an installable progressive web app, so volunteers add it to their home screen and it opens without browser chrome. The capture screen uses the rear camera or the photo library and takes one to six photos per item. The design system has a name, "Verdict on a Shelf Tag." Prices render in a serif display face on a die-cut shelf tag with a chamfered corner and a punched string hole. Status colors stay inside the ministry's rust, gold, and green palette. Violet has exactly one meaning, a price that a person changed before approving it. Staff asked for that so they could spot human-adjusted prices at a glance. Each recommendation carries a research level. Green is a normal result. Yellow leaves the call to the volunteer and puts the checklist up front. Red always goes to manager review. Managers work a review queue where they can approve, set a different price, return the item with a note, or mark it as not priceable. Admins set the review threshold, which defaults to $75, and decide whether electronics, jewelry, and collectibles always need review. Admins also create accounts, since there is no self-service sign-up, and can archive and restore items. The usage page measures participation instead of throughput. My first version credited only the person who set the final price, which hid most of the team. Now everyone who added, priced, or reviewed an item shows up. Solomon is one Next.js 15 app deployed to AWS with SST 4 and OpenNext, running on Lambda with response streaming behind CloudFront. I've written about why SST fits this kind of app. Cost drove most of Solomon's infrastructure choices. At the store's volume, Lambda, S3, CloudFront, and the database all sit inside AWS free tiers. The recurring costs are model calls, a few cents per priced item, and the completed-sales data plan. Photos go straight from the phone to a private S3 bucket through presigned upload URLs. The signature covers the content type, so the bucket only accepts the image formats the server approved. When the item is saved, the server copies each upload under the item's own prefix. A lifecycle rule deletes anything left in the uploads prefix after seven days, which cleans up abandoned captures without a cleanup job. Display and model calls use short-lived presigned reads. No photo is ever public. Authentication started on Clerk and moved to Better Auth on the second day of the project. Better Auth is free and self-hosted, and its tables live in the app's own database, which removed the last outside vendor. Sign-in is email and password. Admins create users and reset passwords from an in-app page, so there is no email provider to configure. The database is Aurora DSQL. Its free tier covers this workload, and IAM-signed connections mean there is no database password to rotate. It is Postgres-compatible, but it leaves out several things I would normally reach for. It has no foreign keys and restricts ALTER TABLE. Each transaction can hold one DDL statement, and index builds run asynchronously. Those limits pushed integrity into the application. Category, condition, and status columns are plain text validated by Zod at the boundary instead of enums or CHECK constraints, because adding a category later would otherwise mean rebuilding a table. Repositories enforce referential integrity. Migrations run through a small custom runner that executes one statement per transaction, rewrites index creation to the async form, and waits for the build job to finish. That wait is a stored procedure, so it needs CALL. I found that out when SELECT raised an error during a migration. There is no local DSQL emulator, so the integration suite runs against a real development cluster. Without foreign keys, a few other rules carry more weight. Item status changes are conditional updates that only succeed if the item is still in the expected state, so two people tapping Accept at the same moment cannot price an item twice. Every state change writes its audit row in the same transaction. Deleting an item is a soft archive, which keeps audit links working. Money is stored as integer cents everywhere, and dollars exist only in the UI and at the model boundary. All dates go through one helper pinned to America/New_York. Lambda runs in UTC, and before that helper existed a history filter for July 7 started at 8 p.m. Eastern on July 6. Early pricing runs used web search with medium reasoning. They usually took 30 to 60 seconds and could spike well past that. By early August, p95 latency on the live stage was about 152 seconds. In SST, the Next.js server.timeout setting drives two values, the Lambda timeout and the CloudFront origin read timeout. The account's CloudFront quota caps the read timeout at 120 seconds, so raising server.timeout past that would push CloudFront over its quota. The read timeout measures the gap between bytes, though, not the total response time. The pricing route streams server-sent events. It sends a start event, a heartbeat every five seconds, and exactly one result. A transform raises the Lambda alone to 300 seconds, and server.timeout stays at 120. The result is a timeout ladder where each limit sits inside the next one: Two other rules make it safe on flaky phone connections. The database write commits before the result event, so a client that sees the result can trust a refresh. A dropped connection never aborts the run. If a volunteer's phone locks mid-run, the recommendation still saves, and the pricing panel polls the item's status until it lands. A "start over" option appears only after the lock expires. A related problem showed up at sign-in. After a quiet spell, the first login paid for a Lambda cold start and a database wake-up and felt broken. One warm instance fixed the container but not the database, because the warmer event never reached a route. A cron now makes a real request through CloudFront to a health endpoint every four minutes, which keeps the whole path warm. Most of the engineering went into a question that sounds easy. What is this item worth, and how do we know? Each of these failures came from real items, and each one left a guard in the code. A volunteer priced a framed cross-stitch map. Solomon came back with three "strong" sold comparables, a framed mallard duck print, a Home Sweet Home sampler, and a daily-bread sampler. The model had matched on medium and framing and missed the subject entirely. It also looked like the app had changed history. The old run's comparison links didn't seem to match what it had claimed, which suggested data corruption. The database was fine. The app had never captured the listing title and ID when it found the evidence, so nobody could see afterward what the model had compared against. The fix had two parts. Solomon started capturing each listing's title, stable item ID, canonical URL, and retrieval time when the evidence arrived, and it never rewrote them on later runs. Comparables that matched only on medium or framing counted as context and could not set a price. Before that change merged, I ran the real judgment prompt against sanitized copies of the original candidates, and it classified all five wrong-subject listings as context only. Every run still stores the exact listings the model was shown. That extra step mattered because of what the regression test proved. A review of the branch showed the test pre-labeled the duck print and samplers as different subjects. It proved the policy rejected bad labels. It didn't prove the model would produce those labels, and a confidently wrong model could have admitted the same evidence. A volunteer priced a stand mixer, then added a photo of the model label with a note saying the model number was updated. Solomon read the label, corrected the model, replaced all three comparables, recalculated the market range, noticed a cracked pouring shield, and kept the price at $99.99 pending a speed and gearbox test. That was a defensible answer. To the volunteer, it looked like nothing had happened. The model hadn't ignored anything. The interface never showed what changed, so the same price read as a non-answer. I also found a real bug. On follow-up results, every checklist item rendered as complete, including checks nobody had done. Follow-ups now produce structured findings, separated into accepted and unverified. The server computes identity, price, range, and evidence changes, and the card says whether Solomon raised, lowered, or kept the price. Checklist items have stable IDs and resolve only when evidence resolves them. An end-to-end test replays the mixer scenario. It asserts the same price before and after, the corrected model, and the mechanical checks still open. I didn't want Solomon moving a price just to prove it had listened. The obvious source of truth for used prices is eBay, so at the start of August I designed a direct integration. The plan was a small vision call to identify the item, the eBay Browse API for listings, and sold data from Marketplace Insights if the account could get it. Rollout would move through off, shadow, and on modes, with the old pricing path authoritative until the new one proved itself. It got complicated fast. eBay rejected my first developer registration with a generic message, and a second one was approved the next day. The production keyset was marked non-compliant until it handled Marketplace Account Deletion notifications. Solomon stored listing facts as audit evidence, so the no-data exemption didn't apply, and I built a stage-fenced, rate-limited endpoint for eBay's verification challenge and notifications. Marketplace Insights is a limited-release API whose documentation sits behind a developer sign-in, and developer approval didn't grant access to it. Browse alone returns active listings, which are asking prices, not sales. The first shadow benchmark fell back on all three cases, with mandatory fees unknown. What changed the direction was a product decision, not a technical one. The useful comparison for the store is what similar items actually sold for. Shipping is helpful context when it's known, and buyer fees don't matter. With that contract, I tested SoldComps, a service that returns completed eBay sales. Fifteen probe searches had a 1.6-second median and a 3.4-second p95. It found relevant sold evidence for eight of nine eligible test cases, and five sampled prices matched the underlying eBay sold pages exactly. The pull request that switched to SoldComps removed about 13,600 lines and added about 6,300. OAuth, the Browse client, the deletion callback, its rate limiter, the admin toggle, and shadow mode all went away. SoldComps has trade-offs I accepted on purpose. It is an independent service built on public eBay sold pages, not a licensed eBay feed, and it offers no SLA or accuracy warranty. Its terms allow internal pricing tools and prohibit reselling raw responses or caching them at scale. eBay doesn't disclose the amount of an accepted Best Offer, so on those rows the listed sold price is only a ceiling. Solomon excludes them. When I turned it on, I still wanted written confirmation about keeping listing evidence for audits, and I went ahead knowing that question was open. The hardening plan from the cross-stitch and mixer incidents was large. In early August, before the switch to SoldComps, an AI coding agent implemented it as a single pull request. It held 141 commits across 235 files, with about 75,600 lines added. Source code went from roughly 12,200 lines to 29,100. Much of the work was good. It added immutable provider facts and a clear split between model judgments and server arithmetic. Wrong-subject evidence failed closed, and follow-up outcomes became explicit. All 1,697 unit tests passed. CI still failed, because the generated OpenNext image optimizer was missing the Linux ARM64 Sharp package. The unit test for that artifact built a fake ARM64 bundle, so it passed while the real artifact couldn't load. Passing unit tests only proved that the mocked pieces worked. I didn't merge it and I didn't throw it away. I posted a decomposition plan on the pull request with five child pull requests and a final configuration-only activation step. The children covered the CI and build baseline, the incident and follow-up fixes, the direct eBay pipeline, the account deletion callback, and photo recovery, each with its own gates and dependency order. The first child pinned Sharp and libvips and added a check that loads the native ARM64 packages from the real generated artifact. CI runs that check on every pull request and every push to main. The original pull request is still open as an integration reference. During a manual review of test results, I found a pair of branded work overalls priced from three sold listings. Two of them were the same item relisted under a new ID two days later, with the same seller, title, photos, and $69.99 price. Counting both put the median at $69.99 and the tag at $38.99. Collapsing the duplicate gave a $44.95 median and a $24.99 tag. The suppression rule is conservative. Two listings collapse only when the normalized title, exact sold price, full-resolution image asset, and condition all match and the sales are within 30 days of each other. The newest row wins. The server drops the image signal before anything leaves the provider boundary, so it never reaches the model, logs, or stored evidence. The rule has a cost. Admission also drops rows without a valid eBay image asset, which removed 55 of 535 raw rows in the next test run. One photo of an iPod nano clearly showed "4GB," model A1137, and the EMC number on its back. Apple's identification guide maps A1137 to the first-generation nano. The identity step transcribed all three correctly, then put "iPod nano" in its exclusion terms, generalized to a generic click-wheel iPod, and asked for a power-on test to confirm an identity that was already visible. A fresh run found exact first-generation sales and rejected them as the wrong generation. Two things were wrong. Visible identifiers weren't binding, and uncertainty about whether the item worked was leaking into what the item was. The fix keeps photographed model and capacity evidence through the conservative identity path and removes exclusions that contradict the identity. Untested operation now counts as a condition question. Condition is now required, and the capture form asks for category-aware facts such as function, packaging, and completeness. When the identity isn't certain, the server searches with a broader, defensible name rather than a guessed model. The server redacts identifiers that look like unit serial numbers before any search, prompt, or display. Two integration test runs were accidentally pointed at the stage the store uses. They left 16 "Test User" items at the top of the history page, four of them with broken thumbnails pointing at fake photo keys. The real photos were fine. Soft archive turned out to be the cleanup tool. I archived exactly those 16 rows and confirmed the real history and thumbnails were intact. Then I added a guard to the integration test configuration that refuses the store's stages before any test or database code loads. It also refuses to run when the stage metadata is missing or malformed. The SoldComps path was strict. It required two defensible sales and dropped anything with a material mismatch. On a fixed 20-case live test set, it priced 4 of 18 eligible items directly, and after the relist fix, 3 of 18. Failed attempts usually spent 30 to 40 seconds before falling back to the old web-search method. Most of those fallbacks were legitimate. There were too few matching sales, or condition, completeness, and variant gaps ruled the matches out. I accepted that coverage and latency as monitored rollout metrics rather than inventing an activation threshold I couldn't justify. An empty SoldComps key worked as the kill switch, and I rehearsed that rollback on a development stage before turning it on. Then the first busy day came in. It brought twelve real items, most of them entered in a little over an hour, and every one went to review. Eleven were low-confidence fallbacks. Median model time was 55 seconds. Staff changed seven of the eleven completed tags, and the final prices totaled 17% less than Solomon's suggestions. Retrieval often found 14 to 38 candidates, and the comparability judgment then admitted none of them. Volunteers were using the app, but the evidence pipeline wasn't doing much of the pricing. Around then I stepped back and restated what the project was for. The goal was to replace a quick photo-in-a-chatbot workflow that people found fast and useful. It was never to produce forensic marketplace evidence for every donated item. I had overcorrected. By late September a price came from a chain. First a vision call identified the item and SoldComps returned candidates. Then a price-blind comparability check ran on a smaller, faster model, and a server tier formula required two defensible sales. Anything that fell short went to the older web-search method. Two pricing methods with different behavior ran side by side, and the tier formula's multipliers had no measured link to what this store charges. I started with an over-engineering audit and a strict rule of zero behavior change. That pass removed completed plan documents, single-implementation interfaces, unused UI variants, and tests that only pinned source text or config literals. More than 20,000 lines came out, most of them completed plan documents. Then I built a backtest. It re-prices every staff-decided item with the real pricing runtime, inside one read-only database transaction, and writes nothing to the database or S3. Store calibration anchors are rebuilt for each item using only decisions made before that item was originally priced, so the replay can't see the future. Per-item artifacts contain item text and listings, so they go to a locked-down directory outside the repository. The fairer comparison is the 116 items where staff changed the price. Staff saw the production suggestion before deciding, so accepted prices lean toward production. The built code had to pass four acceptance gates on error, bias, coverage, and latency. The first build missed three of them. Tuning moved the model calls to a priority processing tier, at twice the token price, and raised the SoldComps budget from 15 to 20 seconds. The prompt didn't change. After tuning, all four gates passed:
  • Median absolute error on staff-changed items was 30.0%, against 33.4% for production. The gate was 33% or less.
  • Median bias on staff-changed items was 0%, against +30% for production. The gate was within 10%.
  • The model cited at least one sale on 73.4% of all items. The gate was 70% or more.
  • Total pricing p95 was 38.3 seconds. The gate was 45 seconds or less.
Within 20% of the staff price, the new design landed 29.3% of staff-changed items against production's 22.6%. Median total pricing time was 24.3 seconds. Coverage is the model's own claim that a listing is comparable, not verified comparability. After the identity and search steps, every run makes one pricing call with no tools. It sees up to six photos, the volunteer's facts, a server-established identity, any follow-up notes, at most eight sold listings, and recent decisions from the same store. Each listing appears as an index, title, condition, sold date, and sold price. IDs and URLs are never included. The model cites the listings it relied on by index and says whether each one is the same item or the same kind of item. Then it sets the shelf price. The server owns everything around that call. It drops citations that point outside the list, repeat, or have an invalid relation. It derives the market range from the cited sold prices and applies the store's price endings and review rules. Then it saves the listings shown, the citations, and the dropped count on the audit row. A run with no valid citation still gets a price, but the server marks it as a best-judgment estimate and sends it to manager review. Only a failure of the pricing call itself is terminal. That pull request added about 4,100 lines and removed about 18,600. TypeScript source went from about 17,200 lines to 12,300. It went live in early October. Before trusting any follow-up idea, I measured the noise. Two re-pricings of identical inputs moved about 18 items better and 18 worse by chance. Most of my ideas didn't clear that bar. Letting the model cite looser "same kind" sales raised coverage from 72% to 86%, but accuracy dropped, and newly cited items got worse. Searching with the model's raw query improved 10 items and hurt 19. Removing store anchors, choosing anchors by title similarity, and switching to medium reasoning all landed within noise, and medium reasoning added 1.8 seconds. Taking the median of four pricing runs didn't help. A 35-second SoldComps budget improved coverage a little, left error unchanged, and pushed p95 to 50 seconds, past the gate. One change did ship. A second SoldComps search runs in parallel on a brand-free product-type phrase from a tiny model call, and the results merge before preselection. Against a same-time control, coverage went from 52.2% to 78.8%, best-judgment fallbacks fell from 88 to 39, and provider deadline misses fell from 61 to 6, with accuracy flat. On 13 hand-labeled search misses, the model went from citing no sales to citing sales on 7. The full end-to-end run of that build met the error, bias, and latency gates, with a 26.8% median error on staff-changed items. It missed the coverage gate at 65.2% on a day when SoldComps was slow. A same-day control of the unchanged design scored 33.4% error, against 30.0% the day before. Day-to-day variance is as large as any effect I measured, and 116 items is a small sample. Those caveats are part of the result. SoldComps returns current sales, so most backtest listings postdate the original pricing, and one staff label per item is not a human noise floor. The new pipeline is simpler because the model now makes a judgment the old pipeline tried to make with code. That only works because the guards stayed. Provider admission, relist suppression, identity binding, and server-owned arithmetic survived the rewrite, and citation checks joined them. What staff see changed in a way I like better. Employees see only the sales the model cited. Managers see every listing the model was shown, with each citation labeled as the model's comparability claim rather than a verified match. Solomon records what it was shown and what it relied on, then puts a person in front of the uncertain cases.
  • Next.js 15, React 19, TypeScript, Tailwind CSS 4, and shadcn components on Radix
  • SST 4 and OpenNext on AWS Lambda with response streaming, CloudFront, and Cloudflare DNS
  • Aurora DSQL with Drizzle ORM and a custom migration runner, plus S3 presigned uploads
  • Better Auth 1.6 with employee, manager, and admin roles
  • OpenAI Responses API with strict structured output, validated with Zod 4
  • SoldComps completed-sale data, searched in parallel
  • Vitest with 398 unit tests, a real-DSQL integration suite, Playwright tests against a mock model server, and GitHub Actions CI that verifies the ARM64 build artifact
  • A read-only pricing backtest over real staff decisions
Solomon is still in use at the store. The next step is to re-run the backtest as staff decisions accumulate, since 116 changed prices is a thin basis for category-level conclusions, and to watch settlement telemetry for coverage on slow provider days. The backtest also changed how I work on this app. Before it existed, deleting code meant trusting my instincts. Now a deletion has to clear the same gates as everything else.

Related projects

Obsidian Vault RAG

Open-source retrieval for Obsidian vaults, with local search, a shared MCP service, project-scoped profiles, and source-checked citations for AI agents.

Rollo Label Printer

Personal label-printing tool that finds the real shipping label inside messy return PDFs and images, normalizes it to 4x6, and prints through direct IPP.

Garmin MCP Server

Open-source MCP server that exposes 34 Garmin Connect health and fitness tools to any AI assistant, with parallel fetching and Docker support.

blakemccarn.dev

Portfolio and blog built with Next.js 16, deployed to AWS via SST v4 with Cloudflare DNS, staging environments, and full IaC.

Charlotte Wire & Cable

Full-stack business website for a specialty wire distributor: searchable catalog, admin dashboard, and serverless AWS infrastructure via SST.

Paperless OCR Enhanced

Open-source Paperless-ngx companion service that uses LLM vision models to repair weak OCR and improve downstream document search.

Paperless Knowledge Graph

Document intelligence system that combines Paperless-ngx, Neo4j, pgvector, Strands Agents, cited answers, and an interactive graph UI.

RapidEPR

Founded and built an AI SaaS product that helps service members across five military services write evaluations, performance statements, and award narratives.