CSC379 Technical Writing

Technical WritingUnit 715 min read

Incident Reports & Problem-Solving Docs: Structure, Writing & Real-World Use

Unit 7 of Technical Writing covers how to write clear, structured incident reports and problem-solving documentation for IT, business, and technical contexts—including definitions, key components, real-world applications (e.g., Ncell outages, eSewa payment failures), and step-by-step templates for reports, memos, and t

Key Concepts & Definitions

1. What is an Incident Report?

An incident report is a formal document that records:

  • What happened (the incident),
  • When and where it occurred,
  • Who was involved (users, systems, or personnel),
  • The impact (e.g., downtime, data loss, security breaches),
  • Root cause and corrective actions.

Purpose:

  • Accountability: Tracks responsibility for errors or failures.
  • Prevention: Helps organizations avoid future incidents.
  • Compliance: Meets legal/industry standards (e.g., ITIL for IT services).
  • Learning: Documents lessons for training or process improvements.

2. What is Problem-Solving Documentation?

This includes:

  • Troubleshooting guides (step-by-step fixes for technical issues).
  • Root cause analysis (RCA) reports (detailed breakdowns of why a problem occurred).
  • Post-mortems (retrospective analyses after major failures, e.g., system crashes).

Key Difference:

Incident Report Problem-Solving Documentation
Focuses on recording an event. Focuses on solving or preventing it.
Written immediately after the incident. Written after analysis (may take days).
Audience: Managers, compliance officers. Audience: Engineers, developers, teams.
Example: "Server X crashed at 3 PM." Example: "How to fix Server X’s memory leak."

Structure of an Incident Report

A well-structured report follows this 5-step framework (adapted from ITIL and ISO standards):

**INCIDENT REPORT**
**Company/Organization**: [Name]
**Date of Incident**: [DD/MM/YYYY, Time]
**Reported By**: [Your Name, Role]
**Incident ID**: [Unique Code, e.g., INC-2024-001]

---
### 1. **Incident Summary** (1-2 sentences)
Briefly describe the incident in **plain language** (avoid jargon).
*Example*:
> *"On 15 May 2024 at 14:30, the eSewa payment gateway froze for 2 hours, preventing 500+ users from completing transactions during peak hours."*

---
### 2. **Details of the Incident**
Use the **5Ws and 1H** framework:
- **Who**: Affected users/systems (e.g., "All Ncell prepaid users in Kathmandu").
- **What**: Exact issue (e.g., "API timeout error in payment processing").
- **When**: Start/end time, duration.
- **Where**: Location (e.g., "eSewa server farm in Lalitpur").
- **Why (Initial Cause)**: Observed symptoms (e.g., "High CPU load due to unoptimized queries").
- **How**: Steps to reproduce (e.g., "Users trying to pay > Rs. 5000 triggered the error").

**Pro Tip**: Include **screenshots/logs** (attach as appendices) if technical.

---
### 3. **Impact Assessment**
Quantify the damage using **SMART metrics** (Specific, Measurable, Achievable):
- **Financial**: "Lost Rs. 250,000 in failed transactions."
- **Operational**: "Customer support calls increased by 300%."
- **Reputational**: "Negative media coverage on social media."
- **Security**: "Potential data exposure if unpatched."

---
### 4. **Root Cause Analysis (RCA)**
Use the **5 Whys Technique** or **Fishbone Diagram** (Ishikawa) to drill down:
*Example for eSewa Failure*:
1. **Why** did the payment gateway freeze? → *Because the database queries timed out.*
2. **Why** did queries time out? → *Because the cache was full.*
3. **Why** was the cache full? → *Because the auto-cleanup script failed.*
4. **Why** did the script fail? → *Because it lacked admin permissions.*
5. **Why** were permissions missing? → *Because the deployment team skipped the access review step.*

**Tools for RCA**:
- **Technical**: Log analysis (e.g., Splunk), error codes.
- **Non-technical**: Surveys (e.g., "Did users see error messages?").

---
### 5. **Corrective Actions & Follow-Up**
List **immediate fixes** and **long-term solutions**:
| **Action**               | **Responsible Person** | **Deadline** | **Status**       |
|--------------------------|------------------------|--------------|------------------|
| Restart failed services. | DevOps Team            | 15 May 2024  | ✅ Completed     |
| Patch database queries.  | Backend Developers     | 18 May 2024  | ⏳ In Progress   |
| Update access controls.  | Security Team          | 20 May 2024  | ❌ Pending       |

**Include**:
- **Preventive measures** (e.g., "Add monitoring alerts for cache size").
- **Training needs** (e.g., "Retrain deployment team on permission checks").

---
### **6. Appendices (Optional but Recommended)**
- **Logs/Screenshots**: Technical evidence.
- **User Testimonials**: Quotes from affected parties (e.g., "I lost my bus ticket money!").
- **Diagrams**: Network topology if the issue was infrastructure-related.

---

### **Real-World Examples**
#### **1. Ncell Network Outage (2023)**
**Incident**: Ncell’s 4G network crashed in Kathmandu for 6 hours during a football match.
**Report Components**:
- **Summary**: "6-hour 4G outage in Kathmandu Valley on 12 Nov 2023."
- **Impact**: 2M users affected; Rs. 5M in lost revenue.
- **RCA**: Faulty fiber optic cable at Thapathali junction (physical damage).
- **Fix**: Bypassed the cable; long-term: redundant fiber routes installed.

**Why It Matters**:
Ncell’s incident report was used to **negotiate compensation** with users and **improve disaster recovery plans**.

#### **2. eSewa Payment Gateway Freeze**
**Scenario**: During Dashain, eSewa’s system failed for high-value transactions (>Rs. 5000).
**Problem-Solving Doc**:
- **Troubleshooting Guide**: "If API returns `504 Gateway Timeout`, check Redis cache."
- **Post-Mortem**: "Scaled database vertically; added load balancers."

**Real Fix**:
eSewa’s team wrote a **public blog post** explaining the issue (transparency builds trust).

#### **3. Daraz Order Fulfillment Delay**
**Incident**: 10% of orders in Pokhara were delayed due to warehouse misrouting.
**Report**:
- **Root Cause**: GPS error in delivery partner app.
- **Solution**: Implemented **geofencing** to verify delivery zones.

---

### **Worked Example: NTC Power Outage Report**
**Situation**: Nepal Electricity Authority (NEA) reports a blackout in Birgunj.
**Your Role**: IT Support at a local hospital using backup generators.

```markdown
**INCIDENT REPORT**
**Organization**: Birgunj Hospital
**Date**: 10 June 2024, 15:47–18:30
**Reported By**: Rajesh KC, IT Coordinator

---
### 1. Incident Summary
NTC’s power outage disrupted hospital operations for 2.5 hours, forcing reliance on backup generators (limited to critical systems like ICU and surgery).

---
### 2. Details
- **Who**: All non-critical departments (admin, billing, labs).
- **What**: Unscheduled power cut; backup generators failed after 1.5 hours.
- **When**: 15:47–18:30 (peak patient admission time).
- **Where**: Birgunj Hospital, Ward 3.
- **Why (Initial)**: NTC’s transformer tripped due to overloading (confirmed via NTC hotline).
- **How**: Staff manually switched to backup, but generator fuel ran low.

---
### 3. Impact
| **Area**          | **Impact**                                  |
|-------------------|---------------------------------------------|
| Patients          | 12 emergency cases delayed.                 |
| Revenue           | Rs. 45,000 lost in unprocessed bills.       |
| Reputation        | Complaints on social media (#NTCFail).     |

---
### 4. Root Cause
1. **Why** did generators fail? → *Fuel tank empty.*
2. **Why** was fuel low? → *Maintenance team forgot to refill (shift change at 15:00).*
3. **Why** was this overlooked? → *No automated fuel-level alerts.*

---
### 5. Corrective Actions
| **Action**                          | **Responsible**       | **Deadline** | **Status** |
|-------------------------------------|-----------------------|--------------|------------|
| Refill generator fuel.              | Maintenance Team     | 10 June      | ✅ Done    |
| Install fuel-level sensors.         | IT Team               | 15 June      | ⏳ Pending |
| Train staff on backup protocols.    | Hospital Admin        | 20 June      | ❌ Pending |

---
### 6. Appendices
- **Screenshot**: NTC outage tweet confirming transformer issue.
- **Log**: Generator fuel log (shows last refill on 9 June).

Problem-Solving Documentation: Troubleshooting Guide

Example: Fixing a "Blue Screen of Death" (BSOD) in Windows.

**TROUBLESHOOTING GUIDE: Windows BSOD (IRQL_NOT_LESS_OR_EQUAL)**
**Audience**: IT Support Staff
**Last Updated**: 2024-05-20

---
### **Step 1: Identify the Error**
1. Note the **STOP code** (e.g., `0x0000000A`).
2. Check **Event Viewer** (`Win + X` > Event Viewer > Windows Logs > System).

---
### **Step 2: Common Causes**
| **Cause**               | **Symptoms**                          | **Solution**                                  |
|-------------------------|---------------------------------------|-----------------------------------------------|
| Corrupt driver          | BSOD after installing new software.    | Roll back driver or update via Device Manager. |
| RAM failure             | Random crashes, memory dump files.    | Run `memtest86` (bootable USB).               |
| Overheating             | Loud fans, shutdowns.                 | Clean dust; reapply thermal paste.           |
| Malware                 | Unusual processes in Task Manager.    | Scan with Windows Defender.                  |

---
### **Step 3: Step-by-Step Fix**
**For Driver Issues**:
1. Press `Win + X` > **Device Manager**.
2. Expand **Display adapters** or **Network adapters**.
3. Right-click the device > **Properties** > **Driver** tab.
4. Click **Roll Back Driver** (if available) or **Update Driver**.

**For RAM Issues**:
1. Download [MemTest86](https://www.memtest86.com/).
2. Create a bootable USB and run for 8 passes.
3. If errors appear, replace the RAM stick.

---
### **Step 4: Prevention**
- **Update drivers** monthly.
- **Monitor temperatures** with tools like HWMonitor.
- **Enable BSOD logging** (`System Properties` > Advanced > Startup and Recovery > Write debugging info).

Comparison: Incident Report vs. Memo

Feature Incident Report Memo (Internal Communication)
Purpose Formal record for compliance/analysis. Quick internal update (e.g., to a team).
Audience Managers, compliance officers, legal. Colleagues, supervisors.
Tone Neutral, detailed. Concise, action-oriented.
Structure 5Ws + RCA + Appendices. TO/FROM, Subject, Key Points, Action Items.
Example Use Case Reporting a data breach to the IT director. Telling the dev team about a server issue.
flowchart TD
    A["Incident Report"] -->|"Focus"| B["Recording an event"]
    A -->|"Audience"| C["Managers, Compliance"]
    A -->|"Timing"| D["Written immediately after"]
    A -->|"Structure"| E["5Ws + 1H + RCA"]

    F["Memo"] -->|"Focus"| G["Internal communication"]
    F -->|"Audience"| H["Team members"]
    F -->|"Timing"| I["Written anytime"]
    F -->|"Structure"| J["Brief, no RCA"]
Key differences between an incident report and a memo (visualized)

Memo Example (for the BSOD issue):

**TO**: DevOps Team
**FROM**: IT Support Lead
**DATE**: 20 May 2024
**SUBJECT**: Urgent: BSOD Reports in Workstations

**Issue**: 5 workstations in Floor 3 crashed today with `IRQL_NOT_LESS_OR_EQUAL`. Suspect corrupt NVIDIA driver (updated yesterday).

**Action Needed**:
1. Roll back drivers on affected machines (see attached screenshot).
2. Test with `memtest86` if crashes persist.
3. Report findings to me by EOD.

**Deadline**: Resolve by 22 May.

**Attachments**: Event Viewer logs, list of affected PCs.

Advantages & Challenges

Advantages of Incident Reports

  • Legal Protection: Proves due diligence (e.g., if sued for downtime).
  • Process Improvement: Identifies recurring issues (e.g., NTC’s frequent outages).
  • Team Accountability: Clarifies who owns fixes.

Challenges

Challenge Solution
Bias in RCA Use data (logs, user feedback), not opinions.
Over-documentation Stick to the 5W1H framework; avoid fluff.
Blame culture Focus on systems, not individuals.

Example of Bad vs. Good RCA:

  • ❌ "The junior dev caused the crash." (Blame)
  • ✅ "The CI/CD pipeline lacked a rollback step for failed deployments." (Systemic)

Exam Tip

How This Unit is Tested

  1. Structure Questions (30%):

    • Expect prompts like "Describe the components of an incident report" or "How would you write a troubleshooting guide for a printer error?"
    • Do: Use the 5W1H + RCA template. Label sections clearly.
    • Don’t: Skip the impact assessment (examiners love quantifiable data).
  2. Scenario-Based Writing (40%):

    • You’ll get a real-world problem (e.g., "Write an incident report for a Daraz delivery delay").
    • Steps to Score Full Marks:
      1. Read carefully: Note the audience (e.g., Daraz’s logistics manager vs. a customer).
      2. Use a template: Start with Summary, then Details, then Impact.
      3. Add real data: Even if hypothetical, invent metrics (e.g., "150 orders delayed").
      4. Propose solutions: Show you understand the problem (e.g., "Improve GPS tracking").
  3. Problem-Solving Docs (20%):

    • For troubleshooting guides, use:
      • Bullet points for steps.
      • Tables for cause-solution pairs.
    • For post-mortems, include:
      • Timeline (e.g., "Issue detected at 10 AM, fixed at 2 PM").
      • Lessons learned (e.g., "Add automated alerts").
  4. Common Pitfalls:

    • ❌ Vague language: "The system broke." → ✅ "The payment API returned a 500 error due to a null pointer exception."
    • ❌ No appendices: Always attach screenshots/logs if the question implies technical issues.
    • ❌ Ignoring the audience: A customer-facing report needs empathy; a manager’s report needs data.

Model Answer Starter (for Exam)

Question: "Write an incident report for a power outage at your college lab that disrupted a programming exam."

**INCIDENT REPORT**
**Organization**: XYZ College, Computer Science Lab
**Date**: 18 April 2024, 11:30–13:00
**Reported By**: Anil Shrestha, Lab In-Charge

---
### 1. Incident Summary
Unscheduled power outage during the **CSC201 Exam** disrupted 45 students for 1.5 hours, delaying the exam by 45 minutes.

---
### 2. Details
- **Who**: 45 students (Batch B), 2 proctors.
- **What**: Complete power loss; UPS backup lasted 30 minutes.
- **When**: 11:30–13:00 (exam started at 11:00).
- **Where**: Lab 3, Ground Floor.
- **Why (Initial)**: NTC’s transformer failure (confirmed via hotline).
- **How**: Students saved work manually; proctors used mobile flashlights.

---
### 3. Impact
| **Area**          | **Impact**                                  |
|-------------------|---------------------------------------------|
| **Academic**      | 45 students lost 45 minutes of exam time.   |
| **Reputation**    | Complaints to the principal’s office.       |
| **Logistics**     | Extra invigilation required for makeup exam.|

---
### 4. Root Cause
1. **Why** did the UPS fail early? → *Battery age (5 years old).*
2. **Why** was this not replaced? → *Budget approval pending.*

---
### 5. Corrective Actions
| **Action**                          | **Responsible**       | **Deadline** | **Status** |
|-------------------------------------|-----------------------|--------------|------------|
| Schedule makeup exam.               | Exam Coordinator     | 20 April     | ✅ Done    |
| Replace UPS batteries.              | Maintenance Team     | 30 April     | ⏳ Pending |
| Train staff on backup protocols.    | Lab In-Charge        | 25 April     | ❌ Pending |

---
### 6. Appendices
- **Screenshot**: NTC outage tweet.
- **Log**: UPS battery health report (shows 60% capacity).

Final Checklist for Full Marks

  1. Structure: Follow the 5W1H + RCA format.
  2. Clarity: Use bullet points and tables for readability.
  3. Data: Include metrics (time lost, users affected).
  4. Solutions: Propose both immediate and long-term fixes.
  5. Professionalism: Avoid emotional language; stick to facts.

Based on the TU BSc CSIT syllabus for Technical Writing (CSC379), unit 7.

Discussion

Loading…