Foundation First! - sitereliability.in (write) - ratsniper.online ( fun) - DevOps | SRE | Cloud | System Design

Seems like some of the A,A+ companies are my dream companies, that make me below average DevOps engineer. its time to go more and more aggressive learning. Image Credits: (propeers)
1
72
I’m locking in for the next 90 days. 12 years of Linux, Datacenter , DevOps, SRE, Cloud & Architecture — I’m going back through all of it and refreshing everything from the ground up. I’m going to share the whole journey here — notes, learnings, experiments, mistakes, and everything in between. If you’re into infra, cloud, DevOps, SRE, or just love tech, keep an eye on this profile. This is going to be one hell of a journey - That’s my promise.
1
3
209
Did anyone start using software factory?
2
189
Someone builds something cool, and suddenly everyone wants to copy it - even the domain name. There are so many “I Hate PDF” clones all over the internet now. I hate this. Build something new, clones . Or at least pick a different domain, yaar #ihatepdf
1
5
184
Replying to @pranvtwt
lots of ihate bro, I hate this.
1
1
28
It's time to sleep 🙃
127
How LY Corp. Runs 1,300+ Kubernetes Clusters (40,000+ nodes ) on Private Infrastructure!!!! Just read this CNCF case study on LY Corp. The problem they had: - Provisioning environments often took more than a week - Releases required cross-team coordination - Nodes remained unpatched for extended periods, increasing security risk - Recovering failed nodes could take several days - Maintenance and upgrades risked service interruption at scale What they did: The team designed a custom resource called KubernetesCluster to represent the desired state of each cluster. Custom controllers continuously reconcile that state on the company’s OpenStack-based private cloud by: - Provisioning virtual machines - Configuring load balancers - Bootstrapping control plane components - Creating and scaling worker nodes Provisioning a new cluster became as simple as applying a resource definition. The result: - 15 engineers operating at hyperscale - Provisioning reduced from weeks to hours - Zero manual node recovery - 100% lifecycle-aligned node rotation every 3–4 months I'll attach the full article link in comment, Please go through it. Let's learn (and unlearn) together #devops
1
125
I’ve been using Cloudflare for a few hobby projects over the last couple of weeks, and I really like how it has evolved, especially the generous free tier for experimenting and building prototypes. It’s no longer just a traditional CDN. Cloudflare is turning into a full-fledged cloud platform, while keeping the developer experience surprisingly simple. - Workers → deploy serverless APIs - Pages → host frontend websites - D1 → edge SQLite database - R2 → object storage / buckets - VPC → private networking - AI Workers - is moving beyond “run a serverless function” toward running AI agents at the edge. - MCP server & Wrangler CLI → honestly, one of my favorite parts I’ve connected Wrangler with Antigravity coding agent for direct deployments, which makes the whole experience really smooth. If you’re building a prototype and don’t want to deal with unnecessary cloud complexity, Cloudflare is a really nice, modern, developer-friendly option. I know there are better dev friendly platforms like Railway,Vercel, Fly etc- but nothing beat Cloudflare's free tier. Of course, it has limitations compared with mainstream cloud platforms like AWS or Azure But — it’s all about the trade-offs. :-)
185
Do you know why your #AWS ECS services could take 3-5+ minutes to scale under a traffic spike? It wasn't necessarily your scaling policy — i think metric resolution could be a major bottleneck. With the default 60-second ECS service metrics, scaling decisions were based on less frequent data, which could delay detection and the subsequent scale-out action. AWS just introduced ECS auto scaling to 20-second high-res metrics, that leads to our scale-out trigger time dropped 363s → 86s (4.2x faster). - As per benchmark. If you've set your scaling threshold low just to buy time for slow detection, you may be able to tighten that threshold now that trigger time dropped 4.2x. #DevOps #ECS #SRE #interviewquestions #systemdesign
2
169
Do you know Why #Cloudflare Workers can start in milliseconds while AWS Lambda can have higher cold-start latency? → AWS Lambda: runs your function inside a secure, isolated execution environment. On a cold start, Lambda prepares the environment, loads the function code and layers, initializes extensions, bootstraps the language runtime, and runs your function's initialization code before invoking the handler. → Cloudflare Workers: run JavaScript inside lightweight V8 isolates. Instead of starting a separate language runtime for each Worker, Cloudflare can run hundreds or thousands of isolates inside the same process. Each isolate gets its own memory and execution context, while the underlying V8 runtime is shared. This removes much of the startup overhead associated with launching a new process/runtime. Cloudflare says isolates can start around 100× faster than a Node.js process in a container or VM. V8 isolate: lightweight language-level sandbox → much lower startup/resource overhead. And this is why the architecture is interesting: Cloudflare isn't removing isolation to gain speed; it's using a much lighter isolation mechanism. It’s not just optimization. It’s a different execution model. (I've attached the cloudflare blog in comment, go through it if you intrested ) #devopsdiaries #devops #sre #systemdesign #devopsinterviewes
3
530
Whoever is in situations, get some free startup credits and host your startup 1-2 year free, I'll share the link in comment. #devops #startup #devopscoach #credits
2
4
177
#DevOps_Diaries: You Can Run "AWS" Entirely on Your Laptop — Free ! Some of you guys probably heard of LocalStack for running AWS locally. But today I want to tell you guys about something better — a tool that actually solves the problem LocalStack created for itself. Its "Floci" For those who don't know LocalStack: it's a tool that simulates AWS services (S3, DynamoDB, Lambda, etc.) on your own machine. Instead of creating real S3 buckets or Lambda functions in AWS and getting billed, you point your AWS CLI/SDK at localhost instead of the real AWS endpoint, and it behaves the same way ( it support same api endpoint so you app can call it like real aws endpoint) — so you can build and test without touching your actual AWS account or bill. Why it's interesting: → No account, no auth token, no feature gates — literally docker compose up and you're running → Startup time: ~24ms vs LocalStack's ~3.3s (yes, I double-checked that number) → Idle memory ~13 MiB vs ~143 MiB → 68 AWS services supported Note: I haven't run this against a heavy production-scale CI pipeline yet, just local testing so far — so if anyone's already swapped a real pipeline over, would love to hear how the actual runtime savings played out at scale. Anyone here still stuck on LocalStack's paid tier for CI? Curious if this closes the gap for you. Let's learn (and unlearn) together 😊 #AWS #DevOps #Docker #Testcontainers #CloudComputing #devopsinterview #devopscoach
4
203
#DevOps Diaries #134 Just read the AWS Architecture Blog on Artera’s prostate cancer AI solution. They take 8GB whole-slide biopsy images, split them into tens of thousands of patches, and run computer vision + multimodal AI on EKS - delivering results in 1–2 days instead of 6 weeks of current available solutions, while keeping the tissue intact for further testing, interesting to see AI is solving reallife pain. Tech Stack: 1. EKS for scalable ML inference/training 2. ECS for clinician web portal 3. S3 + EFS for storage & data locality 4. Global Accelerator sits in front of ALB — routes traffic over AWS global network to cut latency for large file uploads/downloads and improve global availability (HIPAA compliant). Clean example of EKS + global networking for life-critical AI workloads. Full blog link in comment- #AWS #EKS #GlobalAccelerator #Kubernetes #DevOps #AI
1
167
#DevOps Diaries #130 If I had to recommend ONE single AWS service you should learn first… my answer is always IAM. Without proper IAM, nothing else in AWS really matters. I was recently talking to a friend who strongly argued for VPC first. His point was is okay as VPC is important. But here’s the reality: You can build and run services without a custom VPC( lots of managed services ). But without IAM? You’re completely stuck. You can’t create resources, you can’t assign permissions, you can’t secure anything, and every single service becomes useless. IAM is the foundation of AWS security and access control. Everything else sits on top of it. So my strong advice: Don’t just “learn” IAM ==> Be a Master in IAM Understand: - Users, Groups, Roles & Policies - IAM Identity Center (SSO) - IAM Access Analyzer - Least Privilege principle - Policy evaluation logic - Resource-based vs Identity-based policies Trust me : the time you invest in mastering IAM will save you hundreds of hours of debugging weird permission errors later. What’s your #1 AWS service you recommend learning first? Drop it in the comments As always, the comment section is yours - correct me / drop your favourite devops stories, Let’s learn together. #AWS #Cloud #IAM #devopsinterviews #security
3
245
#DevOps Pulse #115 Cloudflare quietly dropped a new Browser Rendering /crawl endpoint yesterday. With a single POST request to api.cloudflare.com/.../crawl you can start from a URL and it will: • Auto-discover the entire site (sitemaps + internal links) • Render pages using a headless browser • Return structured output in **HTML, Markdown, or clean JSON It runs as an async job, so you trigger it once and poll for results later. This basically replaces the usual stack of: Lambda + headless browser + crawler + queue + storage. Super useful for: • RAG pipelines • Website indexing • Automated documentation extraction • AI training datasets DevOps engineers no longer need to build a spider from scratch. #Cloudflare #DevOps #SRE #RAG #Automation
178
#DevOps Diaries #126 Do you know in AWS we can secure items like encryption keys, financial data, or personal information in a way that even the server itself can’t leak them? That’s what AWS Nitro Enclaves do. Imagine creating a tiny, super-secure vault inside your EC2 instance -> completely isolated, with no network, no SSH access, and no persistent storage. Your application can send sensitive data into this #enclave, it gets processed safely, and only the results come back. It’s like having a locked room inside your server that even admins or hackers can’t enter - "a perfect solution for confidential computing." Your daily dose of Byte-Sized Learning—catch you in the next one! :-) #DevOps #AWS #ECS #ECR #SRE
4
328
#DevOps Diaries This is how I’m working on new project now - better ideas are always welcome.! In a new repository, the first question I ask Cursor is to create a /docs folder and act as a teammate. Then we brainstorm together and continuously populate the documentation including mermaid diagrams. Once the brainstorming is complete and the entire documentation is ready, I ask him to switch sides and ask to act as a external consultant to review the entire plan and optimise it once - discussion continuous , once we both are in same page then will build the entire plan - works very well so far 🙃 - earlier I used to build things steps by steps now all in one. #cursor #ai #DevOps
1
1
4
231
Replying to @AskYoshik

ALT Donald Duck Sleep GIF

1
30
#Minio - Dead!!! I like that guy!! but ya its time to say good bye!! #SeaweedFS is pretty mature alternative. #devops #cloud #storage #s3
1
2
390
DevOps Diaries: I just explored AWS newly released "ECS express mode" It is basically a faster, simpler way to run containers on Fargate - you don’t need to configure load balancers, capacity providers, or a lot of ECS service settings; AWS handles most defaults for you. ( it free, no additional cost :-) ) It’s great when you just want to deploy a small API, internal service, or event-driven workload quickly with less infra code and operational overhead. The trade-off is you lose advanced controls like blue-green deployments and complex traffic routing, I think may be in near future they will include those options but still its nice feature, it make our life easier. try it. #AWS #DevOps #containers #SRE
2
347
DevOps Diaries: Today I was debugging a dev-support script on my local and noticed my cursor agent running a command with "rehash" First time I’d seen it , so I read about it - just want to understand what the heck is this rehash? Interesting, your shell doesn’t scan $PATH every time you run a command. That would be too slow. Instead, shells like zsh keep a lookup table (a hash map) of command → full executable path. There is a “command cache” means the shell doesn’t repeatedly search directories in $PATH. So why rehash? Because when you install new tools (like pip adding binaries to ~/Library/Python/3.9/bin), the shell’s lookup table is out of date. Running: export PATH="$HOME/Library/Python/3.9/bin:$PATH" && rehash tells the shell ==> Refresh your command cache — new executables have arrived. 😃 interesting is n't it. #linux #DevOps #zsh #SRE
1
2
346
DevOps Diaries: Do you know why Kubernetes, by default, doesn’t allow or enable swap memory? Swap is one of the greatest features of an operating system, right? So why does Kubernetes “hate” it? I just read an interesting GitHub discussion on this topic. There were a lot of details about why swap wasn’t allowed by default as part of their design choice, why it wasn’t prioritized earlier, and what are the latest changes related to swap usecase etc- thought to share with you guys also - Once you get time please seach and read the thread its worth of your time #AWS #DevOps #Kubernetes #SRE
10
6
144
13,446
DevOps Diaries: Today I just explored an interesting tool, just sharing with you guys also. #LocalStack - This is an open-source tool that lets you run AWS cloud services locally on your machine, without needing an actual AWS account or paying for AWS usage. Instead of deploying your Lambda, S3 bucket, or DynamoDB table to real AWS (which takes time and money), you can run them locally using LocalStack - all in a Docker container. Community version covers most basic AWS services, but there are limitations - but enough to get started and see how it works. This is how I created a S3 locally and upload files: 1. aws s3 mb s3://bucket1 --endpoint-url=localhost.localstack.cloud:4… 2. aws s3 cp j s3://bucket1 --endpoint-url=localhost.localstack.cloud:4… upload: ./j to s3://bucket1/j 3. aws s3 ls s3://bucket1 --endpoint-url=localhost.localstack.cloud:4… 2025-11-12 08:32:03     0 j This is fun, is n't it? 😃 Note: I’m not sure how often they update their APIs to keep up with AWS changes. There are also some limitations when using it for real development work. If anyone has tried replacing their AWS dev playground with LocalStack, Please share your experience. #AWS #DevOps #CloudComputing #SRE
11
63
431
27,097
DevOps Daily: 2 Mins Only 23% of people answered this correctly - it's the "file name" that is not part of the Linux inode inode (index node) is a data structure in Linux that stores all metadata about a file, except its name. It contains details like: 1. File type – Regular file, directory, symbolic link, block device, etc. 2. Permissions 3. Owner, Group 4. File size 5. Timestamps (atime,mtime,ctime) 6. Link count 7. Pointers to data blocks 8. File flags Etc- Etc- - Each file has a unique inode number within its filesystem, inode is the file’s identity & filename is just a label pointing to it. - As always, the comment section is yours:- share your insights, point out corrections, or drop related article links. Let’s learn (and unlearn) together. 🙂 #linux #kernel #devops #sre #devopsinterview
1
14
663
90% of DevOps/SRE engineers (including me at first 😅) ended up choosing the wrong answer in my recent poll - I told you it was a tricky one! The question was simple: (output of below command in bash) - name="abc" && echo $name{1..3}.txt Most people expected the output to be: - abc1.txt abc2.txt abc3.txt But the actual output is: - .txt .txt .txt Why? 🤔 It all comes down to how Bash expands variables vs. braces: 1. Brace expansion happens before variable expansion. So , - {1..3} is expanded first → 1 2 3. 2. After expansion, the command effectively becomes: - echo $name1.txt $name2.txt $name3.txt But variables like $name1, $name2, $name3 are not defined. So they just vanish → leaving only .txt .txt .txt. Next time you see unexpected output in Bash, remember, t’s not always about syntax, it’s about the order of expansions. Bash expansion order really matters: - Brace expansion > Tilde expansion > Parameter expansion > Command substitution > Arithmetic expansion > Word splitting > Globbing #DevOps #SRE #Linux #InterviewPreparation
5
42
1,240
AWS Lambda internals are very interesting, Do you want to know How it working under the hood? Check out my medium post - medium.com/aws-in-plain-engl… #lambda #aws #devops #sre #devopsinterview
1
6
457
Replying to @GhostCoder_
ThinkPADDDD 🫶 GUM GUM!!!!!!!!!!!!!!!!!!!!!!!!
1
22
Great Article - How to Talk Technical Stuff with a Non-Technical Audience? newsletter.systemdesigncodex…
1
187
↕️ HOW #DNS Resolve IP? A Simple Illustration 🖥️ #DevOps
3
327
🚢"In Kubernetes, CPU request sets guaranteed resources for a container. CPU limit caps max usage, preventing disruption. Requests aid scheduling, limits prevent resource abuse. Both balance stable operation and cluster efficiency."🖥️ #DevOps #kubernetes Credits:Natan Yellin 🩷
1
6
411