DevOps Workflow: Deploy safely and frequently
Learn deployment patterns, CI/CD setup, and how to ship features with confidence while maintaining system stability.
DevOps Workflow: Deploy Safely and Frequently
The DevOps Mindset
Old way: Developers write code. Operations deploys it. When something breaks, they blame each other.
DevOps way: Development and Operations collaborate. Developers feel responsibility for production. Operations enables deployment, not blocks it.
Recommended Deployment Architecture
Continuous Integration (CI)
Every code change triggers an automated pipeline:
Developer pushes code
↓
Automated tests run
↓
Build check (does it compile?)
↓
Linting and security scan
↓
Pass? Merge to main
Fail? Reject, notify developer
Benefits:
- Bugs caught before humans review
- Team always works with passing code
- No manual "build" step (error-prone)
Continuous Delivery (CD)
Code that passes CI is ready to ship but requires human approval.
Merged code
↓
Deploy to staging environment
↓
PMs/QA verify in staging
↓
Approve for production
↓
One-click deploy to prod
Continuous Deployment (Extra Challenging)
Every merged code deploys to production automatically. Requires:
- Excellent test coverage (>90%)
- Sophisticated feature flags
- Real-time monitoring
- Instant rollback capability
Most teams should aim for Continuous Delivery, not Deployment.
Safe Deployment Patterns
1. Blue-Green Deployment
Keep two identical production environments running:
┌─────────────┐ ┌─────────────┐
│ BLUE │ │ GREEN │
│ (Current) │────→│ (New) │
│ Users here │ │ Deploy here │
└─────────────┘ └─────────────┘
All good? Switch traffic to GREEN.
Problem? Switch back to BLUE instantly.
Advantages:
- Zero downtime
- Instant rollback
- Easy A/B testing
2. Canary Release
Roll out to a small % of users first:
Day 1: 1% of users see new feature
↓ [Monitor for errors]
Day 2: 10% of users see new feature
↓ [Monitor for errors]
Day 3: 50% of users
↓ [Monitor for errors]
Day 4: 100% of users
If error rate increases at any step, halt rollout and investigate.
3. Feature Flags (Feature Toggles)
Deploy code to production but hide behind a flag:
if (featureFlags.newCheckoutFlow) {
showNewCheckout();
} else {
showOldCheckout();
}
Usage:
- Deploy Friday afternoon (reduces risk)
- Keep feature hidden until Monday
- Flip flag when ready
- Disable instantly if issues arise
Database Migrations
Database changes are the scariest part of deployments. Approach:
Expand, then Contract
Phase 1: Expand (Backward Compatible)
- Add new column/table
- Old code still works
- New code starts using new field
- Deploy once old code is gone
Example:
-- Add new column
ALTER TABLE users ADD COLUMN email_verified BOOLEAN DEFAULT false;
-- Old code: SELECT * FROM users (still works)
-- New code: SELECT *, email_verified FROM users (uses new field)
-- After all code using new column ships...
-- Remove old column (if applicable)
Test Migrations
- Always test migration forward AND backward
- Test on production-scale data (not tiny test db)
- Have rollback plan documented
Production Monitoring & Alerting
Key Metrics to Watch
Availability
- Is the service up? (check every 30 seconds)
- Alert if down for > 1 minute
Performance
- API response time (p95: 200ms, p99: 1s)
- Database query time (p95: 50ms)
- Error rate (alert if > 0.1%)
Business
- Transactions per second
- Checkout success rate
- User signups
Infrastructure
- Disk usage (alert if > 80%)
- Memory usage (alert if > 85%)
- CPU (alert if sustained > 70%)
Alerting Philosophy
Too many alerts = alert fatigue = ignored alerts = missed real issues
Rule: Every alert should be actionable and urgent.
❌ Bad alert: "Memory is 75%" (maybe it's fine, maybe it's not)
✅ Good alert: "Memory > 85% for 5 minutes — restart app?" (clear action)
Incident Response
Before Incident (Prevention)
- Monitoring in place
- Runbooks written (step-by-step fix procedures)
- Escalation path clear (who do I call?)
- Team trained on common failures
During Incident
- Declare incident — Slack channel + video call
- Gather information — What broke? When? What changed?
- Triage severity:
- P1: Production down, customers affected → all hands on deck
- P2: Degraded (slow/errors) → senior engineer + on-call
- P3: Low impact → resolve in next sprint
- Mitigate fast — Don't fix perfectly, fix quickly
- Communicate — "We're aware, investigating, ETA 30 min"
After Incident (Post-Mortem)
- Write what happened (timeline + root cause)
- Discuss as a team — "Why wasn't this caught?"
- Identify improvements (add monitoring? better tests? automate?)
- Create tasks for improvements
- Track: incidents per month should decrease over time
Your First CI/CD Pipeline
Minimal Setup
- Code repository (GitHub, GitLab)
- Automated tests run on every PR
- Staging environment — anyone can deploy by merging to
stagingbranch - Production environment — merge to
main+ click button = deploy
Tools
- CI/CD: GitHub Actions (free, easy)
- Hosting: AWS / Azure / Heroku (your choice)
- Monitoring: Sentry (errors) + DataDog (metrics)
- Databases: Managed (AWS RDS, Azure SQL) not self-hosted
Deployment Checklist
Before you hit "deploy":
- All tests passing
- Code reviewed and approved
- QA sign-off (if applicable)
- Runbook updated
- On-call engineer available to monitor
- Database migrations tested
- Feature flags configured
- Monitoring alerts armed
Common DevOps Mistakes
❌ Deploying on Friday afternoon without monitoring coverage
❌ No rollback plan — "Umm, how do I undo this?"
❌ Database changes without migration scripts — Manual work, error-prone
❌ No monitoring — "App is fine... oh wait, it's been down for 2 hours"
❌ All-or-nothing deploys — Big changes = high risk
Next Steps This Week
- Set up CI/CD for your repository
- Add automated tests to your pipeline
- Deploy to staging
- Document your runbooks
- Set up basic monitoring (error tracking)
Goal: Ship 10x faster, with more confidence.
