Skip to content
Satyam Sharma
Page progress
← All work
Live

Satyam Sharma — MoneyDock

Visit

Solo engineer — architecture, build, data pipelines, operationsPersonal-Finance Platform · FinTech

MoneyDock is a free, no-signup personal-finance platform for Indian retail investors, live at moneydock.in. It does four things: runs 27 financial calculations that recompute live as you drag a slider; serves a page per NSE-listed stock and per mutual fund with computed returns, annualised volatility and maximum drawdown; compares any two of them head-to-head with a ₹10,000-growth overlay; and reads the market — indices, forex, commodities, crypto, and gold and silver rates with the India duty and GST maths made explicit.

I built and operate all of it: the Next.js application, the MongoDB schema, the five ingestion pipelines, the 272-test suite, the SEO architecture that took it to roughly 9,250 indexed pages, and the decisions about what not to build. There is no team, and there is no separate ops function.

The interesting engineering here is not the calculators — it is that this is a money-adjacent product at programmatic scale. Both halves of that phrase constrain the design hard. Money-adjacent means a wrong number is worse than no number. Programmatic scale means thousands of near-identical pages, which is precisely the pattern Google’s March 2026 core update was built to demote. Most of what follows is the consequence of taking both of those seriously at the same time.

168KGoogle search impressions · 27 Nov 2025 – 10 Oct 2026
1,270organic clicks · same window
~9,250pages indexed by Google (4 Oct 2026) · 15,977 in the sitemap
272automated tests · a failure blocks the deploy
27financial calculators, hand-derived
5live external data feeds
Architecture

01 / 09

Problem

Indian retail investing has an information problem that is not a lack of information. Calculator sites exist in abundance; most gate results behind a signup, bury them under ad interstitials, or quietly get the compounding convention wrong — recurring deposits compounded monthly instead of quarterly, income tax computed against last year’s slabs, HRA exemption applying two of the three statutory tests instead of all three.

The security-level data is worse. Trailing returns are everywhere. Annualised volatility and maximum drawdown — the two numbers that actually describe what holding something feels like — are published for almost no direct mutual-fund plans, because computing them requires several years of daily NAV history and nobody wants to store it.

And the obvious way to build at this scale is the way that now gets a site removed from the index. Templating 2,300 stock pages that differ only by name and price is textbook scaled-content abuse. So the problem was never “can I generate thousands of pages” — it was whether each page could contain something that exists nowhere else, computed from data I actually hold.

02 / 09

Goal

Correctness goal

No number on the site may be wrong, and no number may be invented. Where a value cannot be sourced, the page shows fewer facts rather than a plausible guess. Two consequences I held to: P/B ratio and dividend yield were left off stock pages entirely because there is no feed for them, and every metric function returns null rather than NaN so a missing input renders as “N/A” instead of as a number.

Differentiation goal

Every programmatic page must pass one test before it ships:

  • Does this page contain computed numbers that exist nowhere else on the internet in this combination?
  • If the answer is no, the page set does not get built — regardless of how much traffic the keyword has.
  • One concept was killed on exactly this test: per-stock brokerage and STT pages. Charges are a function of trade value and broker, not of the stock, so 2,300 pages would have differed only by name.

Operational goal

One person must be able to run this. That ruled out anything requiring a human in the loop — no editorial approval queue, no per-page LLM generation to review, no manual data entry. Everything either computes deterministically or does not exist.

Performance goal

No network round-trip for a calculation, and no language model anywhere in the request path. Both are TTFB problems and crawl-budget problems before they are user-experience problems.

03 / 09

Architecture

Six layers, arranged so that the expensive and failure-prone work happens on a schedule rather than during a request. Nothing a user waits for talks to an external service.

  1. Ingestion — five external feeds, on cron

    NSE Bhavcopy (the official end-of-day CSV of every listed security), NSE’s equity master list, AMFI’s daily NAV file, Yahoo Finance for intraday and forex, and CoinGecko. Nine cron route handlers, each Bearer-gated on a shared secret, each holding itself to a 45-second time budget and writing through bulkWrite in unordered 500–1,000-operation batches. Every outbound call has an AbortSignal timeout — twelve of them, from 8 seconds for IndexNow to 30 for AMFI — because a hung fetch in a serverless function is billed until it is killed.

  2. Persistence — MongoDB, time series as parallel arrays

    Nine Mongoose models. Price and NAV history are not one document per day; they are compact parallel arrays on a single document per entity, capped server-side with $push … $slice: -1300, which is about five years of trading days. That is ~150 MB for 2,400 stocks and ~90 MB for 3,500 funds, versus tens of millions of documents the obvious way. The connection is a singleton cached on globalThis so a warm serverless instance reuses it.

  3. Computation — pure, null-safe, unit-tested

    Everything financially load-bearing lives in I/O-free modules: trailing returns on a 248-trading-day year, annualised volatility as the standard deviation of daily returns scaled by √248, maximum drawdown, a SIP backtest against real NAV history, and XIRR. Every function ends in a Number.isFinite check and returns null rather than propagating NaN. Feed parsers are split into a pure parse(text) half and a fetching half so the parsers are testable without a network.

  4. Rendering — ISR, tuned per data cadence

    Server Components throughout, with revalidation windows matched to how fast the underlying data actually changes: 600 seconds for the home page and market views, 6 hours for index constituents (Bhavcopy is end-of-day), 12 hours for individual stock and fund pages, 24 hours for blog posts and comparisons. generateStaticParams pre-builds the highest-traffic pages; the long tail renders on demand and is then cached.

  5. Delivery — CDN cache, deterministic fallbacks

    Read APIs set s-maxage with stale-while-revalidate so repeat traffic is absorbed at the edge rather than at the database. Recharts is behind four dynamic imports with ssr:false to keep it out of the initial bundle. Every read endpoint returns the exact shape the client already renders on failure — {a:null,b:null}, or an empty array — so an upstream outage degrades to “N/A” in a table instead of a 500.

  6. Operations — the deploy is the gate

    The Vercel build command is `npm run test:deploy && npm run build`. There is no separate CI step to ignore and no badge to go stale: if the suite fails, there is no deployment. Index-gating logic decides per page whether it is worth indexing at all, and the same predicate filters the sitemap, so the two can never disagree.

The design rationale for the programmatic page sets, including the sets I decided against, is written up in the repository as a scaling blueprint rather than living in my head. Same for the hardening pass — 14 numbered findings, each with the failure it closed and the test that pins it.

Screens pending — live product linked above

04 / 09

Key features

  1. 27 financial calculators

    SIP, EMI, income tax under both regimes, PPF, NPS, GST, HRA, gratuity, EPF, FIRE, SWP and more — recomputing live as sliders move

    Why it is built this wayAll client-side state, so there is no network round-trip per keystroke and no server cost per calculation

  2. Per-fund SIP backtest

    Buys at the first available NAV of each calendar month from real AMFI history and solves XIRR on the resulting cash flows

    Why it is built this wayThe site’s most defensible asset — a real backtest, not a formula projection. Five refuse-to-answer gates run before it computes anything

  3. Volatility and drawdown per security

    Annualised volatility (σ × √248) and maximum peak-to-trough drawdown, computed from stored daily history

    Why it is built this wayAlmost nobody publishes these for direct plans. They are the numbers that make each page unique, which is what makes the page set defensible

  4. Head-to-head comparison

    Any two stocks or funds, with a ₹10,000-growth overlay, a risk table and a data-derived verdict. Pages are generated automatically for pairs in related categories

    Why it is built this wayReversed slugs canonicalise to one sorted URL, so a-vs-b and b-vs-a never compete as duplicates in the index

  5. Deterministic verdict prose

    Every sentence of every generated assessment is computed from that page’s own figures

    Why it is built this wayZero hallucination risk, zero per-page model cost, no approval queue — and no generic filler that would be identical across thousands of pages

  6. Embeddable calculators

    Any of the 27 as a one-line iframe carrying a “Powered by MoneyDock” attribution link

    Why it is built this wayThe only link source that scales without outreach. Embed routes are noindex and get frame-ancestors * while every other route is locked to SAMEORIGIN

  7. Bilingual evergreen content

    English and Hindi post pairs, linked by hreflang, drawn from a finite 45-topic library

    Why it is built this wayThe library is finite on purpose — when every topic is covered the job publishes nothing rather than inventing filler

05 / 09

Technical decisions

The six calls below are the ones I would want to be questioned on. Two of them reversed an earlier decision of my own, which is the part I would rather talk about than the architecture diagram.

  1. One CSV instead of 2,300 API calls.

    Pricing the NSE universe symbol-by-symbol through Yahoo got rate-limited, so most stock pages carried stale or missing prices — and a stock page with no price is worthless. NSE publishes an official end-of-day Bhavcopy: one CSV with the closing price of every listed security. Switching to it replaced roughly 2,300 per-symbol requests with a single one. The tradeoff is that the single request becomes load-bearing, so it is defended four ways: dual-host retry across two NSE archive domains, a 15-second timeout, a browser User-Agent and Referer because NSE rejects non-browser agents, a sanity floor that treats a CSV under 100 rows as failure rather than as data, and a ten-day backward walk so weekends and market holidays resolve without intervention.

  2. Removing the language model from the request path.

    The comparison pages originally called Gemini synchronously on the first visit to any new pair. It worked, and it was wrong: multi-second TTFB, a crawl-budget sink, and a request that could hang until the platform killed it into a raw error page. I rebuilt it as a two-speed page. The most recent 500 pairs are pre-built; everything else renders instantly from a deterministic fallback computed from stored figures, marked noindex-but-follow so a thinner version never enters the index yet still passes link equity, and upgraded later by a cron job. When I later built fund comparisons, they never had a model in the request path at all — that is the lesson already learned, applied without having to relearn it.

  3. Deterministic prose over generated prose, deliberately against my own plan.

    The plan had been to generate fund-comparison verdicts with Gemini. I replaced that with computed sentences from unit-tested functions. Three reasons: no hallucination risk on a page about someone’s money, no per-page inference cost across thousands of pairs, and no approval queue — which matters because I am the only person who could staff one. The constraint that made it work is a rule in the assessment module: a bullet is emitted only when the underlying metric exists, so missing data produces fewer bullets, never invented ones.

  4. Fail-open for indexing, fail-closed for authentication.

    The same codebase makes opposite choices about ambiguity, and the direction is chosen by blast radius. The index gate defaults to indexable — anything undefined, NaN, or the wrong type is treated as “index it”, because wrongly hiding a page that is earning impressions costs more than wrongly indexing a thin one. The admin auth module defaults to locked — if the password variable is not configured, no request is ever treated as admin, so a misconfiguration leaves the endpoints shut rather than silently open. It also bails on length before comparing bytes, so a wrong password cannot be discovered one character at a time through response timing.

  5. Time series as capped parallel arrays.

    Five years of daily closes for 2,400 stocks is roughly three million rows the obvious way. Storing them as parallel arrays on one document per entity and capping with $slice: -1300 keeps it at about 150 MB, and reading one entity’s full history is a single document fetch instead of a range scan. The cost is that arbitrary cross-entity queries over history become awkward — I accepted that, because the product never needs one. The chart endpoints then project only the fields they use, which halved the bytes read per request when I noticed the high and low arrays were being fetched and discarded.

  6. Letting robots.txt allow what Googlebot needs to render.

    A blanket disallow on /_next/ looks like sensible hygiene. It meant Googlebot was rendering the site completely unstyled — confirmed by looking at the screenshot in Search Console’s live test, not by reasoning about it. The fix carves /_next/static/ and /_next/image back out, relying on robots.txt longest-match precedence. The same file then explicitly allows seventeen named AI and answer-engine crawlers, on the view that being readable by the systems that increasingly answer these questions is worth more than the traffic it might divert.

06 / 09

Challenges

  1. Scaling without triggering the filter designed to catch it.

    The March 2026 core update demoted template-with-variable-substitution at scale by 50–90%, and programmatic pages are this site’s entire growth thesis. There was no way to test the boundary safely, so I released in gated batches — 500, then 2,000, then the rest — advancing only when Search Console showed at least 60% of a batch indexed, with a hard stop if “crawled, currently not indexed” passed roughly 40%.

  2. XIRR has no closed form, and NAV history has holes.

    Schemes change, holidays interrupt, series start mid-life. A silently wrong return figure on a page about someone’s savings is worse than no figure, so the solver bisects net present value over a bounded range for up to 200 iterations and returns null when no sign change exists rather than reporting a bogus root. Five gates run before it: array-length parity, a minimum of two months, sufficient purchase months, every NAV strictly positive, and exact calendar-month contiguity — computing the expected span and refusing if it disagrees with the observed count, so a gappy series cannot quietly misrepresent returns.

  3. One database blip poisoning the process until redeploy.

    The classic serverless singleton bug, and the hardest class to catch because it is invisible in normal operation. The connect promise was cached on globalThis, so a single failed connection during a cold start cached the rejected promise — and every later request on that warm instance replayed the same rejection until the process happened to restart. Every page on the site returning 500, with no error in the request that caused it. Three lines fixed it; what makes it closed rather than patched is a test that pins both halves, asserting the retry happens and that the happy path still connects exactly once, so a future refactor cannot quietly open a connection per request.

  4. A finance page that must not fabricate.

    The temptation on a thin page is always to add a plausible-looking field. I left P/B and dividend yield off stock pages because there is no source for them. Where a stock has no industry classification, the peer heading says “most actively traded stocks” rather than claiming a relationship that does not exist. This is a discipline problem more than an engineering one, and it is the constraint I found hardest to hold.

  5. Reading my own metrics correctly.

    Mass indexation makes average position look like collapse: new pages enter the index around position 70 and drag the mean down while nothing that was ranking has moved. Pruning genuinely empty pages does the reverse — indexed count dips and average position improves, which looks like a loss and is the intended outcome. Learning to read the two apart took longer than building the pipeline that caused them.

  6. Automated content that was actively harmful.

    An earlier version of the blog job produced roughly 95 posts of news, IPO updates, stock tips and Hindi duplicates — content stale within days, competing against publications this domain cannot outrank, and dragging the whole site under a scaled-content signal. I audited it, cut about 80 posts, and rewrote the job to draw from a finite library of 45 evergreen topics that self-terminates instead of churning. Deleting most of what an automation I wrote had produced was the correct call and not a comfortable one.

07 / 09

Lessons learned

  1. Choose fail-open or fail-closed per blast radius, not per codebase.

    I used to treat defensive defaults as a single house style. Holding two opposite defaults in one codebase — and being able to say why each one points the way it does — turned out to be a much better way to reason about failure than picking a side.

  2. A test suite is worth what it blocks.

    The suite grew from 100 to 272 tests, but the change that mattered was putting it in the build command rather than in CI. A green badge is advisory. A build that will not produce a deployment is a constraint, and I stopped having to remember to be careful.

  3. Lock behaviour before you improve it.

    Before the hardening pass I wrote 100 tests against the unmodified source and got them green, so any fix that changed observable behaviour would be visibly wrong rather than arguably fine. It cost a day up front and made the following fourteen fixes boring, which is exactly what you want from a hardening pass.

  4. Returning null is a design decision.

    Every metric function returning null rather than NaN means the honest answer travels all the way to the page and renders as “N/A”. Once that convention existed, “we don’t have this number” stopped being an error case to suppress and became a normal value the UI knows how to display.

  5. Confirm rendering assumptions by looking.

    The robots.txt problem was invisible to reasoning and obvious in a screenshot. I now check what a crawler actually receives rather than what the configuration ought to imply — the same instinct as reading the response instead of trusting the client.

  6. The best decision on the project was a deletion.

    Killing the per-stock brokerage pages, cutting 80 blog posts, and refusing to add two fields I had no source for did more for this site than any feature I shipped. Restraint is not a substitute for engineering, but at programmatic scale it is part of it.

08 / 09

Future improvements

In roughly the order I would take them, and stated plainly rather than as a roadmap. The first three are gaps I would raise myself in an interview.

  1. No IP or user rate limiting exists. Protection today is structural — CDN caching absorbs repeat traffic, cron routes are Bearer-gated, admin writes are token-gated — but that is not the same thing, and it is the first item I would close.

  2. There is no end-to-end suite. 272 unit and integration tests cover the maths, the parsers, the API contracts and the component states, but nothing drives a real browser through a calculator and asserts the rendered figure.

  3. The income tax and Atal Pension calculators use rounded slab offsets and an approximation formula rather than the exact statutory tables. Both are marked as approximations in the source; on a finance site both should be exact.

  4. One dependency (sharp) is imported by the blog image job but missing from package.json — it resolves only through a transitive copy, which is a build waiting to break.

  5. Only one of the nine cron jobs currently pings IndexNow on publish. Wiring the rest would cut the lag between a page updating and search engines seeing it.

  6. The sitemap is a single document. It is fine at ~16,000 URLs and needs sharding before it reaches the 50,000-URL limit; the rationale for deferring it is written down rather than assumed.

  7. Coverage reporting is not configured, so I know the suite is meaningful by reading it rather than by measuring it.

09 / 09

Technologies used

In the codebase

  • Next.js 16 (App Router)
  • React 19
  • TypeScript (strict)
  • MongoDB
  • Mongoose
  • Tailwind CSS
  • Vitest
  • Testing Library
  • Recharts
  • AWS S3
  • Vercel
  • Google Gemini
  • ISR
  • JSON-LD structured data

Engineering practices the build depends on

  • Time-series modelling
  • ETL pipeline design
  • Financial mathematics
  • XIRR / numerical methods
  • Caching strategy
  • Programmatic SEO
  • Test-gated deployment
  • Failure-mode analysis

Next case study

Animatrixx

Read the case study