"It was a customer messaging me on LINE to say the site wouldn't load — that's how I found out." Before that message arrived, the site might have already been down all night: people opened it, saw an error page, and left; orders got submitted but never actually went through — and you had no idea any of it happened. It's not that you don't care about quality; no one has been watching that server for you in the middle of the night. The following three things are the smallest monitoring setup you need.
The most awkward way to find out: your customer tells you
By the time a customer tells you, the site has usually been down for a while: maybe it went down overnight and nobody noticed until morning, or maybe a feature has been broken for three days and everyone assumed they'd clicked the wrong thing. You have no idea what happened during that gap:
- How many people opened the page, saw an error, and left — never to come back.
- How many orders were submitted but never actually reached the backend.
- How many people quietly decided to switch to a competitor, without you ever knowing.
- How many support calls came in from people whose questions couldn't be answered, because the person who picked up didn't know the system was down either.
What's worse is that even once you know, you don't know where to start looking: is the server down, is the code throwing errors, or is the database full? Without logs, you're just guessing. This is one of the five things covered in A Working Demo Isn't the Same as Production-Ready, and it's also the easiest one to fix first.
The smallest monitoring setup: three things, plus one precondition
"Monitoring" sounds like something only big companies need. In reality, it just answers four questions:
| What it checks | The question it answers | What happens without it |
| Is it alive | Can people open it right now | It goes down overnight and no one notices until morning |
| Is it throwing errors | Is a feature actually broken | Even after a customer reports it, you can't find the cause |
| Server resources | Are disk or memory almost full | It goes down without warning on an otherwise ordinary afternoon |
| Someone receives the alert | Who sees it when something goes wrong | All three are set up, and still nobody responds |
The first three are about tools; the last one is about people. Have all four, and you'll know before your customers do. Have none of them, and you're running on luck and your customers' goodwill.
The first two: is it alive, and is it throwing errors
The first one is the simplest to set up: have something open your site every few minutes and notify you if it can't. The second is a bit harder, because "alive" isn't the same as "working" — the homepage might load fine while checkout does nothing when you click it.
- The check interval needs to be short enough: stretch it too long and a brief outage slips through unnoticed — everything just looks fine.
- The alert needs to actually get your attention: a text message or an instant message both work — sending it to an inbox nobody checks is the same as not having monitoring at all.
- Errors need to be logged somewhere you can see: have the code write down what went wrong, instead of just showing the user an apology.
- A spike in errors needs to trigger an alert: going from three a day to thirty an hour is an incident, not a minor glitch.
Products built with AI tools often skip the second one, because nobody told the AI to "log errors when they happen."
The third thing: can the server keep up
The server is the computer your site runs on, and it has three things that can run out: disk space, memory, and processing power. Run out of any of them, and the site slows down before it goes down. This kind of failure is the most avoidable, because it happens gradually:
- Disk space gets filled up a little at a time by log files and uploaded images — you could have seen it coming a week earlier.
- Memory runs short and the system starts restarting over and over, while users just think, "why is this so slow today."
- Processing power gets stretched thin, usually during a campaign or peak season — exactly the day you can least afford to go down.
- None of the three is full yet, but the trend keeps climbing — this is the cheapest and safest moment to deal with it.
The approach is to check these three numbers regularly and act before they run out. It's the same point made in Does Your AI-Built Product Have Backups: these are all problems that "someone watching" prevents from happening.
You have monitoring — now what: who's watching it
Setting up the three monitoring pieces only solves half the problem. The other half starts after the alert goes out:
- Who receives it: if the alert lands in an inbox nobody checks, the entire setup was pointless.
- Who's able to act on it: whoever gets the message at 2 a.m. needs to understand it and actually be able to fix it.
- How fast they need to respond: without an agreed timeframe, "someone saw it" and "someone fixed it" are two very different things.
- Who writes it up afterward: why it went down this time and how it got fixed — skip the write-up, and it'll happen again.
So monitoring eventually comes back to a more basic question: does this system have someone responsible for it. That's also the conclusion of The Gap Between Vibe Coding and Real System Architecture — AI can help you build the thing, but it won't take responsibility for what happens after it's built.
How Nerdtechnic watches it for you
- We watch the server and act on issues first: not waiting for you to notice, and certainly not waiting for your customers to.
- Response times agreed up front: server or service issues get a response within 1 business day; errors you report get an assessment back within 2 business days.
- Security updates at least once a quarter: urgent cases get handled separately — nothing gets left to pile up until next year.
- Billed monthly, cancel anytime: basic maintenance starts at NT$6,000/month, with no lock-in contract keeping you stuck.
Monitoring isn't just buying a tool — it's finding someone willing to wake up in the middle of the night.
Who should know first when your site goes down? It should be "the person watching it" — not your customer. Nerdtechnic maintains your system as your Helper CTO, functioning like your company's technical lead without being on your payroll. The first step is a 60-minute system health check: we only need to see the code, not your production account credentials, and within 3–5 business days you'll get a one-page report you can actually understand. If your product is still relying on customers as its monitoring system, let us take a look at it as it stands today: Helper CTO: System Maintenance Plan.