You Can Just Choose Not to Have Problems
A lot of problems can be removed by construction instead of energy-consuming piecemeal efforts.
The goodest doggos^Wexample
My partner and I have four dogs. They’re around 40-70 lbs each, so let’s call it 200 lbs of doggo. Our pack has been that size for a number of years, going as high as five and as low as two (a blissfully calm period). They’re mostly old ladies–sort of a staging site for the rainbow bridge situation–and so are fairly relaxed, though one of them is a strapping young pup of a few years now who’s mostly matched the vibe with some training.
Various things we’ve done to deliberately manage the chaos factory:
- No other pets. Other pets mean more things that can go wrong.
- Standardized feeding times and food. A couple of commands gets the dogs in place, they wait, we fill bowls, we dismiss. Everybody eats the same stuff. The rare medical exceptions are a light lift.
- No young dogs. We know our activity level, and new puppies require an amount of care and attention that we’d rather spend elsewhere.
- Close doors to rooms not in use. We have the public area that the dogs inhabit autonomously, and then places like offices or bedrooms are kept closed unless human supervision is present. I work from home–on a day with fires to put out, I don’t also need dog issues in my workspace.
- Kennels and the training to use them. Part of the first firmware updates for the dogs is getting them comfortable being kenneled, overnight if needed. Dogs with known self-control problems get kenneled at night; well-behaved dogs are randomly kenneled to keep training fresh. This means that their chaos potential is contained if we have to go leave the house and can’t have them out–especially useful if a dog has the occasional case of diarrhea and you’d rather solve the problem with a hose and bleach and five minutes instead of a washer and dryer and two hours.
- Scheduled vet visits. Twice a year, half the pack each time, year-long meds (heartworm, etc.).
- Important things kept off dog reach. We keep the floor and low tables basically clear of things interesting to dogs.
- The dogs in public are always leashed. We do not rely on the good training and temperament of our dogs to prevent chases or running off, and we maintain control over them at all times. If we have a problem, it’s usually because some jackass left their leashless but “well-trained” dog to run over and start something.
Each of these is effective not only because of the structure it provides, but the opportunities for chaos it removes.
If a dog can’t get into a room, the dog cannot mess up the room–we don’t have to stay on top of that training (though we do) because the opportunity does not exist. If a dog is kenneled and a lunch runs long, the dog cannot get onto furniture unattended long enough to have an accident–our dogs are house-trained and let out regularly but on the off-chance they wouldn’t be able to hold it the kennel keeps it in one place. If a dog can’t run away from you, the dog is much less likely to get hit by a car or attack another dog.
We have deliberately structured the dog situation to exclude common problems and uncommonly bad outcomes.
As with dogs so too engineers
Engineering teams are, to a first and second approximation, able to benefit from the same approach. Things I’ve done or seen done on various teams:
- No deploys on Friday afternoon. If you don’t deploy on Friday, you don’t run into missing people, grumpy folks postponing weekend vacations, or being unable to reach other key parts of the business.
- Every merge deploys automatically. On the other side of the spectrum, if every merge deploys to prod you almost never accumulate special-case runbooks and you never have to worry about “okay, who deploys this change?” or “okay, which documents do I need to update for the changelog and which channels get notifications about it and do I roll the systemd unit or update symlinks first?”.
- Standup at the same time in the same way every day. If you have a dedicated standup every day, you will drastically reduce the chances of an engineer going feral and being out-of-the-loop for days or weeks. You don’t have to schedule about when the team is going to meet, you don’t have to worry about “oh no Frank is out today we’ll all be syncing again in at least a week from now”, you don’t have to be surprised when somebody has a problem and is unable to find time to bring it up with the person on their critical path–the time is blocked out. You don’t have to worry about sprawling, wending conversations because the format is fixed and terse.
- Work on a feature is logged in the ticket. We’re not going to play silly games about which Slack thread something happened in, what email chain it was in, or whatever else–the work communication is in Jira (or the moral equivalent) and if it isn’t there it didn’t happen. We aren’t going to set ourselves up for having to interrupt engineers when leadership has a question and management needs to do an org stacktrace.
- Automate everything possible. We’re not going to have a spreadsheet where a senior engineer or EM gets tagged once a month to, for a few days, read phone bills and update a Google Spreadsheet and manually back out KPIs and reconcile with other departments–we’re going to automate it and then the manyfold ways of that chain breaking will dry up and cease.
Each of these removes an entire category of ways things can go wrong. Some of them–for example, the standups–are things “free-thinking” engineers will complain about…but the savings in time and mental overhead, if pointed out, usually make a convincing argument after a few weeks.
…and so too software
In software systems, we have the same opportunities! Sufficiently ambitious engineers can just choose not to build software with problems!
- Don’t fragment your systems. You almost certainly don’t need to wrap every single bit of business logic in a microservice, because that will bring in the questions of networking and partitioning and monitoring and how all of that gets pulled together and provisioned.
- Don’t duplicate state across a network with a master-master relationship. One-way flow of state approximates a caching problem, and that’s bearable, but master-master state–say, a thick JS client and a server where they’re both considered authoritative–across a network tends to end up in split-brain problems.
- Don’t support exotic setups unless you have to. There is a user out there using lynx on an ancient SGI O2 to browse your website. You do not need to try and design a system to service that user. It’s okay, I promise. Screen-readers, on the other hand, are an ADA (and just basically decent) thing you should probably look into–and doing it now, in stages, will save you from the annoying problems that show up when legal action is threatened or a client says “hey so we need a11y on here to close the contract” and throws the surprise feature toaster into your planning jacuzzi.
- Don’t pick non-standard languages. Choosing to use snowflake languages instead of boring tech–say, Pony instead of Go, Crystal instead of Ruby, etc.–means that when you inevitably need to debug a problem you will be faced with fewer resources and potentially completely novel problems. Sometimes, a weird language (say, datalog or MATLAB) in a particular problem domain really is the answer and can be unreasonably productive–but if you aren’t sure that that’s your domain it almost certainly isn’t.
- Don’t write your own transport formats. We have JSON. We have msgpack. If you’re a masochist and want footguns, we’ve already invented ASN.1. You don’t want to be figuring out bugs in your transport at the same time as you are figuring out bugs in your application logic.
- Don’t ignore compiler warnings, linter errors, and runtime log spew. The computer is trying to help you, and choosing to ignore the admonitions almost always masks the spread of deeper problems and flaws. Especially in the monitoring on-call state, letting spew pile up means you are guaranteeing more work for yourself to sift through false-positives and noise usually at the most time-crunched and stressful times (e.g., outages).
This is a small sampling of the software world, hopefully enough to get the idea across: you can just choose not to have certain entire categories of problems.
Why do people opt in to problems?
Sometimes, it’s abject laziness. Sometimes, it’s inexperience and ignorance. Sometimes, it’s because they’re rushed past the point where they could be proactive. Sometimes, it’s because they don’t feel the cost of the problems they enable. Sometimes they’re just stupid.
I’ve seen (or been!) every variant of these, and it can be hard to see from the inside when you fall into one of those buckets. It can also be incredibly hard to try and tell others–especially in a team setting–that you think they’re falling into one or more of the buckets and be heard. Practicing kind and direct communication helps.
Held to the fire to propose some unifying theory of opting in to problems, I’d pull on that “not feeling the cost of problems they enable” thread:
Having standardized dog food and using Go is boring–you know you’re giving up fun. Having kennels that you need to use every night and clearing out log spew is a known tediousness. Closing doors in the house and closing old tickets removes potential adventure. In all cases, the cost (however minor) of the boring prophylactic behavior is visible up front, where a more free-wheeling approach is a slot-machine of variable reward and hidden cost. Rust WASM for the input forms? 12 SaaS providers for different parts of the auth flow? A chance to build a resume with Mongo or a columnar store? Twelve interlocking project planning and grooming sessions? Pull the lever, maybe this time it’ll pay out!
The costs of those decisions are usually borne by the rest of the team, or future generations of developers. I get the brief excitement of learning Rust, and months later some poor schmuck who is on-call needs to learn how traits work to trace a problem instead of putting his kids to bed. Sally gets to put on her resume that she knows Kubernetes, and then the business has to pay for AWS consultants to come in and migrate to the next version of EKS because Sally left for a gig at Google after her project was launched on an EOL’ed version of the platform. People opt in to problems that they themselves are unlikely to be inconvenienced by.
You get to decide how much of your dwindling life force you want to expend solving avoidable problems. Do you really want to hang out or work with people that let their dogs and processes off-leash to become problems? Do you want to be that person whose dog or service or clever code bites somebody else who was just trying to make it through the workday?
If you don’t choose, the choice will be made for you–poorly and without regard for your happiness.
Want to discuss this post? | Discuss via email | Hacker News | Lobsters |