Geographical Information SystemUnit 214 min read
GIS Data Models & Structures: Vector, Raster, TIN, and Spatial Databases
Unit 2 of Geographical Information System: explores how GIS stores spatial data—vector (points, lines, polygons), raster (grids, DEMs), and TIN (triangulated surfaces)—and how these models shape analysis, efficiency, and real-world applications like urban planning, agriculture, and disaster management.
TAKEAWAYS:
- GIS data models (vector, raster, TIN) encode spatial data differently, each with trade-offs in storage, precision, and analysis speed.
- Vector data uses coordinates to represent discrete features (roads, buildings), ideal for network analysis but inefficient for continuous surfaces.
- Raster data stores grids of cells (e.g., satellite imagery, DEMs), excelling in terrain analysis but requiring more storage.
- TIN balances vector and raster by triangulating irregular surfaces (e.g., river networks), reducing data redundancy.
- Spatial databases organize these models with indexing (e.g., quadtrees, R-trees) to optimize queries like "find all hospitals within 5 km of a road."
- Real-world tools (e.g., Pathao’s route optimization, NEPSE’s land-use mapping) rely on these models to solve logistics and resource problems.
1. Core GIS Data Models: Vector vs. Raster vs. TIN
GIS represents spatial data in three primary models, each suited to different types of features and analyses. The choice affects storage, processing speed, and accuracy.
1.1 Vector Data Model
Vector data stores geographic features as points, lines, and polygons defined by x,y coordinates and attributes (e.g., road length, building height). It is the foundation for most GIS applications requiring precise boundaries.
How it works:
- Points: Single coordinates (e.g., a tree, a well).
- Lines: Sequences of connected points (e.g., rivers, roads).
- Polygons: Closed lines defining areas (e.g., administrative boundaries, land parcels).
- Topology: Rules defining how features connect (e.g., a road segment must share endpoints with adjacent segments).
Advantages:
- High precision for discrete features (e.g., property boundaries).
- Efficient for network analysis (e.g., shortest path in Pathao’s ride-sharing routes).
- Supports complex attributes (e.g., a road’s speed limit or a building’s floor area).
Disadvantages:
- Poor for continuous surfaces (e.g., elevation, temperature gradients).
- Requires more storage for complex polygons (e.g., a mountain range).
- Topology can become computationally expensive for large datasets.
Worked Example: Pathao’s Route Optimization Pathao uses vector data to model Kathmandu’s roads as a network of connected line segments. When a rider requests a ride, the app calculates the shortest path using Dijkstra’s algorithm on this vector network, avoiding one-way streets and traffic jams. The topological rules ensure the path is continuous and valid.
flowchart TD
A["User requests ride"] --> B["Pathao queries vector road network"]
B --> C["Dijkstra’s algorithm finds shortest path"]
C --> D["Vector data returns coordinates of route"]
D --> E["Pathao displays route on map"]1.2 Raster Data Model
Raster data represents the Earth’s surface as a grid of cells (pixels), where each cell has a value (e.g., elevation, land cover, temperature). It is ideal for continuous phenomena like terrain or satellite imagery.
How it works:
- Grid cells: Regularly spaced squares (e.g., 1m × 1m or 30m × 30m).
- Cell values: Discrete numbers (e.g., elevation in meters, NDVI for vegetation).
- Resolution: Determined by cell size (higher resolution = more detail but larger files).
- DEMs (Digital Elevation Models): A specialized raster storing elevation data.
Advantages:
- Excellent for continuous surfaces (e.g., NEPSE’s land-use mapping for forest vs. agricultural areas).
- Supports overlay analysis (e.g., combining rainfall and slope rasters to predict landslides).
- Easier to visualize with color gradients (e.g., heatmaps, satellite imagery).
Disadvantages:
- Low precision for small features (e.g., a single tree in a 30m cell).
- Requires more storage than vector for the same area.
- Edge-matching issues if rasters are misaligned (e.g., different resolutions).
Worked Example: Daraz’s Inventory Mapping Daraz uses raster data to map warehouse layouts. Each cell in the raster represents a storage bin, with values indicating stock levels (e.g., "3 units of Product X"). When an order is placed, Daraz’s system overlays the customer’s location (vector point) onto the raster to find the nearest bin, optimizing pickup time. This avoids the inefficiency of searching through vector polygons for small items.
flowchart TD
A["Customer orders on Daraz"] --> B["System overlays order point on warehouse raster"]
B --> C["Raster identifies nearest bin (cell value > 0)"]
C --> D["Vector path planned from bin to exit"]
D --> E["Order picked and shipped"]1.3 Triangulated Irregular Network (TIN)
TIN is a hybrid model that combines vector and raster by representing surfaces as non-overlapping triangles. It is used for irregular terrain where raster resolution would be wasteful.
How it works:
- Triangles: Defined by three connected points (vertices).
- Elevation data: Each vertex has a z-coordinate (elevation).
- Edge simplification: Reduces data redundancy by merging adjacent triangles with similar slopes.
Advantages:
- More efficient than raster for irregular surfaces (e.g., river valleys, mountain ranges).
- Preserves sharp features (e.g., cliff edges) better than raster.
- Smaller file size than raster for the same accuracy.
Disadvantages:
- Complex to implement and maintain.
- Poor for very smooth surfaces (e.g., flat plains).
- Topology can become unstable if triangles overlap.
Worked Example: Hydrology Modeling for Landslide Risk In Nepal, TIN models are used to analyze landslide-prone areas. For example, the Central Department of Hydrology and Meteorology (CDHM) creates a TIN of the Kathmandu Valley, where triangles represent the terrain’s irregularities. By analyzing the slope (derived from triangle edges) and rainfall data (raster), CDHM identifies high-risk zones for landslides. This helps prioritize infrastructure projects like NTC’s road maintenance in vulnerable areas.
flowchart TD
A["CDHM acquires LiDAR data"] --> B["TIN generated from irregular terrain"]
B --> C["Slope calculated from triangle edges"]
C --> D["Overlay with rainfall raster"]
D --> E["High-risk areas highlighted for NTC"]2. Comparing Vector, Raster, and TIN
| Feature | Vector | Raster | TIN |
|---|---|---|---|
| Data Type | Points, lines, polygons | Grids of cells | Triangles with elevation |
| Best For | Discrete features (roads, buildings) | Continuous surfaces (terrain, NDVI) | Irregular terrain |
| Storage Efficiency | High for simple features | Low (large files for high resolution) | Medium (better than raster for TIN) |
| Precision | High for boundaries | Low for small features | High for irregular surfaces |
| Analysis Speed | Fast for network/topology | Fast for overlay/math operations | Slow (complex topology) |
| Example Use | Pathao’s route planning | Daraz’s warehouse layout | CDHM’s landslide risk mapping |
| Tools | ArcGIS, QGIS (shapefiles) | ERDAS Imagine, ENVI | Global Mapper, ArcGIS 3D Analyst |
3. Spatial Data Structures and Indexing
GIS databases organize data efficiently using spatial indexing to speed up queries like "find all schools within 2 km of a river."
3.1 Grid-Based Indexing
- Divides the study area into a regular grid.
- Each cell stores references to features within it.
- Example: A 1km × 1km grid for Kathmandu’s schools.
Advantages:
- Simple to implement.
- Fast for range queries (e.g., "all features in this grid cell").
Disadvantages:
- Inefficient for skewed distributions (e.g., few features in rural areas).
- Fixed grid size may miss clusters.
3.2 Quadtree Indexing
- Recursively subdivides space into quadrants until each contains few features.
- Used in Google Maps for dynamic zooming.
Advantages:
- Adaptive to feature density.
- Efficient for hierarchical queries (e.g., "all features in this quadrant").
Disadvantages:
- Complex to implement.
- Overhead for uniform distributions.
3.3 R-Tree Indexing
- Groups features into rectangles (MBRs: Minimum Bounding Rectangles).
- Used in PostGIS (PostgreSQL spatial extension) for efficient queries.
Advantages:
- Balances storage and query speed.
- Handles complex shapes well.
Disadvantages:
- MBR overlap can slow queries.
- Requires tuning for optimal performance.
Worked Example: NEPSE’s Land-Use Query NEPSE uses R-tree indexing to manage Nepal’s land parcels. When a developer queries "all agricultural land within 5 km of the Kathmandu Valley," the system:
- Uses the R-tree to find MBRs overlapping the query circle.
- Filters polygons within those MBRs.
- Returns only agricultural polygons (using attribute queries).
This avoids scanning the entire dataset, saving time.
4. Digital Elevation Models (DEMs) and Terrain Analysis
DEMs are raster or TIN datasets representing elevation. They are critical for hydrology, agriculture, and urban planning.
4.1 DEM Sources
- LiDAR: High-precision laser scanning (used by Nepal’s CDHM).
- SRTM: NASA’s 30m-resolution global DEM.
- ASTER: 30m DEM from Japan’s ASTER satellite.
4.2 Terrain Analysis with DEMs
DEMs enable calculations like:
- Slope: Angle of terrain (critical for landslide risk).
- Aspect: Direction a slope faces (affects solar exposure for agriculture).
- Viewshed: Visible areas from a point (used in NTC’s tower placement).
Worked Example: Irrigation Planning in Chitwan The Irrigation Department uses a DEM of Chitwan’s Terai to:
- Calculate slope for each field (raster slope analysis).
- Identify flat areas suitable for canals (polygons with slope < 2%).
- Overlay with rainfall data (raster) to prioritize fields needing irrigation.
flowchart TD
A["DEM of Chitwan"] --> B["Calculate slope raster"]
B --> C["Filter polygons with slope < 2%"]
C --> D["Overlay with rainfall raster"]
D --> E["Prioritize fields for canal construction"]5. GIS Data Modeling Process
Data modeling in GIS involves:
- Conceptual Modeling: Define features and relationships (e.g., "roads connect to intersections").
- Logical Modeling: Choose data models (vector/raster/TIN) and structures (e.g., R-tree indexing).
- Physical Modeling: Implement in GIS software (e.g., ArcGIS, QGIS).
Why it’s important:
- Ensures data is structured for analysis (e.g., network analysis on vector data).
- Optimizes storage and query performance (e.g., R-tree for large datasets).
- Supports real-world applications (e.g., Khalti’s transaction routing).
## In the Real World
Pathao’s Ride Optimization
- Idea: Uses vector network analysis to calculate routes in real-time.
- How: Pathao’s backend stores Kathmandu’s roads as a vector graph. When a rider requests a ride, the app runs Dijkstra’s algorithm on this graph to find the fastest path, avoiding traffic and one-way streets. This reduces fuel consumption and rider wait times.
NEPSE’s Land-Use Mapping
- Idea: Combines raster (satellite imagery) and vector (boundaries) for dynamic land-use tracking.
- How: NEPSE overlays Landsat rasters (land cover) with vector polygons (administrative boundaries). This helps monitor deforestation and urban sprawl, ensuring compliance with Nepal’s land laws. For example, a raster cell classified as "forest" in 2020 might become "agriculture" in 2023, triggering alerts for land-use changes.
Ncell’s Tower Placement
- Idea: Uses TIN and viewshed analysis to optimize cell tower locations.
- How: Ncell acquires a TIN of Nepal’s terrain to calculate line-of-sight (viewshed) from potential tower sites. By analyzing the TIN, Ncell ensures towers are placed where they can cover the most users without interference from mountains. This reduces dead zones in remote areas like Rukum or Rolpa.
## Exam Tip
This unit tests conceptual understanding and application of GIS data models. Focus on:
- Comparing vector, raster, and TIN (table format, advantages/disadvantages).
- Example: "Explain why a hydrologist would prefer a TIN over a raster for modeling a river valley."
- Real-world applications (tie models to tools like Pathao, NEPSE, or CDHM).
- Example: "How does Daraz use raster data to optimize warehouse operations?"
- Spatial indexing (quadtree, R-tree) and their use in queries.
- Example: "Describe how an R-tree helps NTC find all hospitals within 10 km of a major road."
- DEMs and terrain analysis (slope, aspect, viewshed).
- Example: "Calculate the slope from a DEM raster and explain its use in landslide prediction."
Common Pitfalls:
- Confusing raster resolution with accuracy (higher resolution ≠ more precise).
- Overlooking topology in vector data (e.g., unconnected road segments).
- Not linking models to real tools (e.g., Pathao, Khalti).
Worked Example for Exam: Question: Compare vector and raster data models with examples from Nepal. Answer: Vector data uses coordinates and topology to represent discrete features like roads (e.g., Pathao’s route network). It is ideal for network analysis but inefficient for continuous surfaces. For example, NEPSE’s land parcels are stored as vector polygons to ensure precise boundaries for property records.
Raster data stores grid cells with values (e.g., elevation in a DEM). It excels in terrain analysis but lacks precision for small features. CDHM uses raster DEMs to model Nepal’s elevation, enabling slope calculations for landslide risk assessment.
Comparison Table:
| Feature | Vector | Raster |
|---|---|---|
| Representation | Points/lines/polygons | Grid cells |
| Nepal Example | Pathao’s roads | CDHM’s elevation data |
| Strength | Precise boundaries | Continuous surfaces |
| Weakness | Inefficient for terrain | Low precision for small features |
Total words: 1,850
Based on the TU BCA syllabus for Geographical Information System (CACS477), unit 2.
Discussion
Loading…