Technical WritingUnit 715 min read
Incident Reports & Problem-Solving Docs: Structure, Writing & Real-World Use
Unit 7 of Technical Writing covers how to write clear, structured incident reports and problem-solving documentation for IT, business, and technical contexts—including definitions, key components, real-world applications (e.g., Ncell outages, eSewa payment failures), and step-by-step templates for reports, memos, and t
Key Concepts & Definitions
1. What is an Incident Report?
An incident report is a formal document that records:
- What happened (the incident),
- When and where it occurred,
- Who was involved (users, systems, or personnel),
- The impact (e.g., downtime, data loss, security breaches),
- Root cause and corrective actions.
Purpose:
- Accountability: Tracks responsibility for errors or failures.
- Prevention: Helps organizations avoid future incidents.
- Compliance: Meets legal/industry standards (e.g., ITIL for IT services).
- Learning: Documents lessons for training or process improvements.
2. What is Problem-Solving Documentation?
This includes:
- Troubleshooting guides (step-by-step fixes for technical issues).
- Root cause analysis (RCA) reports (detailed breakdowns of why a problem occurred).
- Post-mortems (retrospective analyses after major failures, e.g., system crashes).
Key Difference:
| Incident Report | Problem-Solving Documentation |
|---|---|
| Focuses on recording an event. | Focuses on solving or preventing it. |
| Written immediately after the incident. | Written after analysis (may take days). |
| Audience: Managers, compliance officers. | Audience: Engineers, developers, teams. |
| Example: "Server X crashed at 3 PM." | Example: "How to fix Server X’s memory leak." |
Structure of an Incident Report
A well-structured report follows this 5-step framework (adapted from ITIL and ISO standards):
**INCIDENT REPORT**
**Company/Organization**: [Name]
**Date of Incident**: [DD/MM/YYYY, Time]
**Reported By**: [Your Name, Role]
**Incident ID**: [Unique Code, e.g., INC-2024-001]
---
### 1. **Incident Summary** (1-2 sentences)
Briefly describe the incident in **plain language** (avoid jargon).
*Example*:
> *"On 15 May 2024 at 14:30, the eSewa payment gateway froze for 2 hours, preventing 500+ users from completing transactions during peak hours."*
---
### 2. **Details of the Incident**
Use the **5Ws and 1H** framework:
- **Who**: Affected users/systems (e.g., "All Ncell prepaid users in Kathmandu").
- **What**: Exact issue (e.g., "API timeout error in payment processing").
- **When**: Start/end time, duration.
- **Where**: Location (e.g., "eSewa server farm in Lalitpur").
- **Why (Initial Cause)**: Observed symptoms (e.g., "High CPU load due to unoptimized queries").
- **How**: Steps to reproduce (e.g., "Users trying to pay > Rs. 5000 triggered the error").
**Pro Tip**: Include **screenshots/logs** (attach as appendices) if technical.
---
### 3. **Impact Assessment**
Quantify the damage using **SMART metrics** (Specific, Measurable, Achievable):
- **Financial**: "Lost Rs. 250,000 in failed transactions."
- **Operational**: "Customer support calls increased by 300%."
- **Reputational**: "Negative media coverage on social media."
- **Security**: "Potential data exposure if unpatched."
---
### 4. **Root Cause Analysis (RCA)**
Use the **5 Whys Technique** or **Fishbone Diagram** (Ishikawa) to drill down:
*Example for eSewa Failure*:
1. **Why** did the payment gateway freeze? → *Because the database queries timed out.*
2. **Why** did queries time out? → *Because the cache was full.*
3. **Why** was the cache full? → *Because the auto-cleanup script failed.*
4. **Why** did the script fail? → *Because it lacked admin permissions.*
5. **Why** were permissions missing? → *Because the deployment team skipped the access review step.*
**Tools for RCA**:
- **Technical**: Log analysis (e.g., Splunk), error codes.
- **Non-technical**: Surveys (e.g., "Did users see error messages?").
---
### 5. **Corrective Actions & Follow-Up**
List **immediate fixes** and **long-term solutions**:
| **Action** | **Responsible Person** | **Deadline** | **Status** |
|--------------------------|------------------------|--------------|------------------|
| Restart failed services. | DevOps Team | 15 May 2024 | ✅ Completed |
| Patch database queries. | Backend Developers | 18 May 2024 | ⏳ In Progress |
| Update access controls. | Security Team | 20 May 2024 | ❌ Pending |
**Include**:
- **Preventive measures** (e.g., "Add monitoring alerts for cache size").
- **Training needs** (e.g., "Retrain deployment team on permission checks").
---
### **6. Appendices (Optional but Recommended)**
- **Logs/Screenshots**: Technical evidence.
- **User Testimonials**: Quotes from affected parties (e.g., "I lost my bus ticket money!").
- **Diagrams**: Network topology if the issue was infrastructure-related.
---
### **Real-World Examples**
#### **1. Ncell Network Outage (2023)**
**Incident**: Ncell’s 4G network crashed in Kathmandu for 6 hours during a football match.
**Report Components**:
- **Summary**: "6-hour 4G outage in Kathmandu Valley on 12 Nov 2023."
- **Impact**: 2M users affected; Rs. 5M in lost revenue.
- **RCA**: Faulty fiber optic cable at Thapathali junction (physical damage).
- **Fix**: Bypassed the cable; long-term: redundant fiber routes installed.
**Why It Matters**:
Ncell’s incident report was used to **negotiate compensation** with users and **improve disaster recovery plans**.
#### **2. eSewa Payment Gateway Freeze**
**Scenario**: During Dashain, eSewa’s system failed for high-value transactions (>Rs. 5000).
**Problem-Solving Doc**:
- **Troubleshooting Guide**: "If API returns `504 Gateway Timeout`, check Redis cache."
- **Post-Mortem**: "Scaled database vertically; added load balancers."
**Real Fix**:
eSewa’s team wrote a **public blog post** explaining the issue (transparency builds trust).
#### **3. Daraz Order Fulfillment Delay**
**Incident**: 10% of orders in Pokhara were delayed due to warehouse misrouting.
**Report**:
- **Root Cause**: GPS error in delivery partner app.
- **Solution**: Implemented **geofencing** to verify delivery zones.
---
### **Worked Example: NTC Power Outage Report**
**Situation**: Nepal Electricity Authority (NEA) reports a blackout in Birgunj.
**Your Role**: IT Support at a local hospital using backup generators.
```markdown
**INCIDENT REPORT**
**Organization**: Birgunj Hospital
**Date**: 10 June 2024, 15:47–18:30
**Reported By**: Rajesh KC, IT Coordinator
---
### 1. Incident Summary
NTC’s power outage disrupted hospital operations for 2.5 hours, forcing reliance on backup generators (limited to critical systems like ICU and surgery).
---
### 2. Details
- **Who**: All non-critical departments (admin, billing, labs).
- **What**: Unscheduled power cut; backup generators failed after 1.5 hours.
- **When**: 15:47–18:30 (peak patient admission time).
- **Where**: Birgunj Hospital, Ward 3.
- **Why (Initial)**: NTC’s transformer tripped due to overloading (confirmed via NTC hotline).
- **How**: Staff manually switched to backup, but generator fuel ran low.
---
### 3. Impact
| **Area** | **Impact** |
|-------------------|---------------------------------------------|
| Patients | 12 emergency cases delayed. |
| Revenue | Rs. 45,000 lost in unprocessed bills. |
| Reputation | Complaints on social media (#NTCFail). |
---
### 4. Root Cause
1. **Why** did generators fail? → *Fuel tank empty.*
2. **Why** was fuel low? → *Maintenance team forgot to refill (shift change at 15:00).*
3. **Why** was this overlooked? → *No automated fuel-level alerts.*
---
### 5. Corrective Actions
| **Action** | **Responsible** | **Deadline** | **Status** |
|-------------------------------------|-----------------------|--------------|------------|
| Refill generator fuel. | Maintenance Team | 10 June | ✅ Done |
| Install fuel-level sensors. | IT Team | 15 June | ⏳ Pending |
| Train staff on backup protocols. | Hospital Admin | 20 June | ❌ Pending |
---
### 6. Appendices
- **Screenshot**: NTC outage tweet confirming transformer issue.
- **Log**: Generator fuel log (shows last refill on 9 June).
Problem-Solving Documentation: Troubleshooting Guide
Example: Fixing a "Blue Screen of Death" (BSOD) in Windows.
**TROUBLESHOOTING GUIDE: Windows BSOD (IRQL_NOT_LESS_OR_EQUAL)**
**Audience**: IT Support Staff
**Last Updated**: 2024-05-20
---
### **Step 1: Identify the Error**
1. Note the **STOP code** (e.g., `0x0000000A`).
2. Check **Event Viewer** (`Win + X` > Event Viewer > Windows Logs > System).
---
### **Step 2: Common Causes**
| **Cause** | **Symptoms** | **Solution** |
|-------------------------|---------------------------------------|-----------------------------------------------|
| Corrupt driver | BSOD after installing new software. | Roll back driver or update via Device Manager. |
| RAM failure | Random crashes, memory dump files. | Run `memtest86` (bootable USB). |
| Overheating | Loud fans, shutdowns. | Clean dust; reapply thermal paste. |
| Malware | Unusual processes in Task Manager. | Scan with Windows Defender. |
---
### **Step 3: Step-by-Step Fix**
**For Driver Issues**:
1. Press `Win + X` > **Device Manager**.
2. Expand **Display adapters** or **Network adapters**.
3. Right-click the device > **Properties** > **Driver** tab.
4. Click **Roll Back Driver** (if available) or **Update Driver**.
**For RAM Issues**:
1. Download [MemTest86](https://www.memtest86.com/).
2. Create a bootable USB and run for 8 passes.
3. If errors appear, replace the RAM stick.
---
### **Step 4: Prevention**
- **Update drivers** monthly.
- **Monitor temperatures** with tools like HWMonitor.
- **Enable BSOD logging** (`System Properties` > Advanced > Startup and Recovery > Write debugging info).
Comparison: Incident Report vs. Memo
| Feature | Incident Report | Memo (Internal Communication) |
|---|---|---|
| Purpose | Formal record for compliance/analysis. | Quick internal update (e.g., to a team). |
| Audience | Managers, compliance officers, legal. | Colleagues, supervisors. |
| Tone | Neutral, detailed. | Concise, action-oriented. |
| Structure | 5Ws + RCA + Appendices. | TO/FROM, Subject, Key Points, Action Items. |
| Example Use Case | Reporting a data breach to the IT director. | Telling the dev team about a server issue. |
flowchart TD
A["Incident Report"] -->|"Focus"| B["Recording an event"]
A -->|"Audience"| C["Managers, Compliance"]
A -->|"Timing"| D["Written immediately after"]
A -->|"Structure"| E["5Ws + 1H + RCA"]
F["Memo"] -->|"Focus"| G["Internal communication"]
F -->|"Audience"| H["Team members"]
F -->|"Timing"| I["Written anytime"]
F -->|"Structure"| J["Brief, no RCA"]Key differences between an incident report and a memo (visualized)Memo Example (for the BSOD issue):
**TO**: DevOps Team
**FROM**: IT Support Lead
**DATE**: 20 May 2024
**SUBJECT**: Urgent: BSOD Reports in Workstations
**Issue**: 5 workstations in Floor 3 crashed today with `IRQL_NOT_LESS_OR_EQUAL`. Suspect corrupt NVIDIA driver (updated yesterday).
**Action Needed**:
1. Roll back drivers on affected machines (see attached screenshot).
2. Test with `memtest86` if crashes persist.
3. Report findings to me by EOD.
**Deadline**: Resolve by 22 May.
**Attachments**: Event Viewer logs, list of affected PCs.
Advantages & Challenges
Advantages of Incident Reports
- Legal Protection: Proves due diligence (e.g., if sued for downtime).
- Process Improvement: Identifies recurring issues (e.g., NTC’s frequent outages).
- Team Accountability: Clarifies who owns fixes.
Challenges
| Challenge | Solution |
|---|---|
| Bias in RCA | Use data (logs, user feedback), not opinions. |
| Over-documentation | Stick to the 5W1H framework; avoid fluff. |
| Blame culture | Focus on systems, not individuals. |
Example of Bad vs. Good RCA:
- ❌ "The junior dev caused the crash." (Blame)
- ✅ "The CI/CD pipeline lacked a rollback step for failed deployments." (Systemic)
Exam Tip
How This Unit is Tested
Structure Questions (30%):
- Expect prompts like "Describe the components of an incident report" or "How would you write a troubleshooting guide for a printer error?"
- Do: Use the 5W1H + RCA template. Label sections clearly.
- Don’t: Skip the impact assessment (examiners love quantifiable data).
Scenario-Based Writing (40%):
- You’ll get a real-world problem (e.g., "Write an incident report for a Daraz delivery delay").
- Steps to Score Full Marks:
- Read carefully: Note the audience (e.g., Daraz’s logistics manager vs. a customer).
- Use a template: Start with Summary, then Details, then Impact.
- Add real data: Even if hypothetical, invent metrics (e.g., "150 orders delayed").
- Propose solutions: Show you understand the problem (e.g., "Improve GPS tracking").
Problem-Solving Docs (20%):
- For troubleshooting guides, use:
- Bullet points for steps.
- Tables for cause-solution pairs.
- For post-mortems, include:
- Timeline (e.g., "Issue detected at 10 AM, fixed at 2 PM").
- Lessons learned (e.g., "Add automated alerts").
- For troubleshooting guides, use:
Common Pitfalls:
- ❌ Vague language: "The system broke." → ✅ "The payment API returned a 500 error due to a null pointer exception."
- ❌ No appendices: Always attach screenshots/logs if the question implies technical issues.
- ❌ Ignoring the audience: A customer-facing report needs empathy; a manager’s report needs data.
Model Answer Starter (for Exam)
Question: "Write an incident report for a power outage at your college lab that disrupted a programming exam."
**INCIDENT REPORT**
**Organization**: XYZ College, Computer Science Lab
**Date**: 18 April 2024, 11:30–13:00
**Reported By**: Anil Shrestha, Lab In-Charge
---
### 1. Incident Summary
Unscheduled power outage during the **CSC201 Exam** disrupted 45 students for 1.5 hours, delaying the exam by 45 minutes.
---
### 2. Details
- **Who**: 45 students (Batch B), 2 proctors.
- **What**: Complete power loss; UPS backup lasted 30 minutes.
- **When**: 11:30–13:00 (exam started at 11:00).
- **Where**: Lab 3, Ground Floor.
- **Why (Initial)**: NTC’s transformer failure (confirmed via hotline).
- **How**: Students saved work manually; proctors used mobile flashlights.
---
### 3. Impact
| **Area** | **Impact** |
|-------------------|---------------------------------------------|
| **Academic** | 45 students lost 45 minutes of exam time. |
| **Reputation** | Complaints to the principal’s office. |
| **Logistics** | Extra invigilation required for makeup exam.|
---
### 4. Root Cause
1. **Why** did the UPS fail early? → *Battery age (5 years old).*
2. **Why** was this not replaced? → *Budget approval pending.*
---
### 5. Corrective Actions
| **Action** | **Responsible** | **Deadline** | **Status** |
|-------------------------------------|-----------------------|--------------|------------|
| Schedule makeup exam. | Exam Coordinator | 20 April | ✅ Done |
| Replace UPS batteries. | Maintenance Team | 30 April | ⏳ Pending |
| Train staff on backup protocols. | Lab In-Charge | 25 April | ❌ Pending |
---
### 6. Appendices
- **Screenshot**: NTC outage tweet.
- **Log**: UPS battery health report (shows 60% capacity).
Final Checklist for Full Marks
- Structure: Follow the 5W1H + RCA format.
- Clarity: Use bullet points and tables for readability.
- Data: Include metrics (time lost, users affected).
- Solutions: Propose both immediate and long-term fixes.
- Professionalism: Avoid emotional language; stick to facts.
Based on the TU BSc CSIT syllabus for Technical Writing (CSC379), unit 7.
Discussion
Loading…