
Building Modern Data Analytics Solutions on AWS
Contact IT Dojo for current pricing, available dates, and a custom quote tailored to your team or organization.
Course Duration
4 Days
Audience
Employees of federal, state and local governments; and businesses working with the government.
Prerequisites
We recommend that attendees of this course have: AWS Technical Essentials Architecting on AWS
Course Description
This intensive four-day bootcamp integrates four sequential AWS courses covering the complete data analytics lifecycle. Students build and secure data lakes with AWS Glue and Lake Formation, run high-performance batch analytics on Amazon EMR with Apache Spark and Hive, design and query Amazon Redshift data warehouses, and build real-time streaming pipelines with Amazon Kinesis and Amazon MSK. Ideal for data engineers, data platform engineers, and solutions architects who need hands-on AWS analytics experience in a compressed timeframe.
Learning Objectives
- Compare data lake and data warehouse architectures and choose the right pattern for a given use case
- Ingest, catalog, and prepare data with AWS Glue, and build data lakes with AWS Lake Formation
- Process large-scale batch workloads on Amazon EMR using Apache Spark and Apache Hive
- Build serverless data processing pipelines with AWS Glue and orchestrate them with Step Functions
- Design, load, and query Amazon Redshift data warehouses, including Redshift Spectrum and semi-structured data
- Build real-time streaming pipelines using Amazon Kinesis Data Streams, Kinesis Data Analytics, and Amazon MSK
- Secure and monitor data lakes, EMR clusters, Redshift warehouses, and streaming pipelines
- Apply modern data architecture patterns across the full analytics lifecycle
Course Outline
Day One: Building Data Lakes on AWS
- Module 1: Introduction to Data Lakes
- Data lake value proposition
- Data lakes vs. data warehouses comparison
- Data lake component architecture
- Common architectural patterns
- Module 2: Data Ingestion, Cataloging, and Preparation
- Storage and ingestion relationships
- AWS Glue crawlers for catalog creation
- Data formatting, partitioning, and compression strategies
- Lab: Building a simple data lake
- Module 3: Data Processing and Analytics
- Data processing applications
- AWS Glue data processing
- Amazon Athena analysis capabilities
- Module 4: Building Data Lakes with AWS Lake Formation
- Lake Formation features and advantages
- Data lake creation procedures
- Security model overview
- Lab: Building a data lake with Lake Formation
- Module 5: Additional Lake Formation Configurations
- Blueprint and workflow automation
- Security and access control implementation
- Record matching with FindMatches
- Amazon QuickSight visualization
- Lab: Blueprint automation
- Lab: QuickSight visualization
- Module 6: Architecture and Review
- Knowledge assessment
- Architecture review
- Course wrap-up
Day Two: Building Batch Data Analytics Solutions on AWS
- Module A: Data Analytics Overview
- Use case examination
- Analytics data pipeline framework
- Module 1: Introduction to Amazon EMR
- EMR in analytics solutions
- Cluster architecture
- Interactive cluster launch demonstration
- Cost optimization strategies
- Module 2: Ingestion and Storage
- Storage optimization techniques
- Data ingestion methods
- Module 3: High-Performance Batch Analytics with Apache Spark
- Spark on EMR use cases and advantages
- Core Spark concepts
- Interactive Spark shell demonstration
- Transformation, processing, and notebook integration
- Lab: Spark low-latency analytics
- Module 4: Batch Processing with Hive
- EMR and Hive integration
- Data transformation approaches
- Introduction to Apache HBase
- Lab: Hive batch processing
- Module 5: Serverless Data Processing
- Serverless transformation and analytics
- AWS Glue and EMR integration
- Lab: Spark orchestration with Step Functions
- Module 6: Security and Monitoring
- EMR cluster security measures
- Client-side encryption demonstration
- Monitoring, troubleshooting, and Spark history review
- Module 7: Batch Analytics Solution Design
- Use case exploration
- Workflow design activity
- Module 8: Modern Data Architectures
- Contemporary architecture patterns
Day Three: Building Data Analytics Solutions Using Amazon Redshift
- Module 1: Redshift in Analytics Pipelines
- Redshift data warehousing rationale
- Platform overview
- Module 2: Amazon Redshift Introduction
- System architecture
- Console tour demonstration
- Feature overview
- Lab: Data loading and querying
- Module 3: Ingestion and Storage
- Ingestion approaches
- Jupyter notebook connection demonstration
- Data distribution strategies
- Semi-structured data analysis with the SUPER data type
- Lab: Redshift Spectrum analytics
- Module 4: Processing and Optimization
- Data transformation techniques and advanced querying strategies
- Lab: Transformation and querying
- Resource management, mixed workload management, and cluster resizing demonstrations
- Module 5: Security and Monitoring
- Cluster security implementation
- Monitoring and troubleshooting procedures
- Module 6: Data Warehouse Solution Design
- Use case review
- Warehouse workflow design activity
- Module 7: Modern Data Architectures
- Contemporary architecture patterns
Day Four: Building Streaming Data Analytics Solutions on AWS
- Module 1: Streaming Services in Analytics
- Streaming data analytics importance
- Streaming pipeline architecture
- Core streaming concepts
- Module 2: Introduction to AWS Streaming Services
- Available streaming services and Kinesis in analytics solutions
- Kinesis Data Streams exploration
- Lab: Kinesis delivery pipeline setup
- Kinesis Data Analytics, Amazon MSK, and Spark Streaming overview
- Module 3: Kinesis for Real-Time Analytics
- Clickstream workload exploration
- Stream creation, producer, and consumer development
- Building and deploying a Flink application
- Lab: Kinesis Data Analytics with Apache Flink
- Module 4: Kinesis Security and Optimization
- Actionable insight optimization
- Security and monitoring practices
- Module 5: Amazon MSK in Streaming Solutions
- MSK use cases and cluster creation procedures
- MSK provisioning demonstration
- Data ingestion, transformation, and processing
- Lab: MSK access control introduction
- Module 6: MSK Security and Optimization
- Optimization strategies and storage scaling demonstration
- Security and monitoring procedures
- Lab: MSK streaming pipeline deployment
- Module 7: Streaming Solution Design
- Use case review
- Streaming workflow design exercise
- Module 8: Modern Data Architectures
- Contemporary architecture patterns
Frequently Asked Questions
What does the Building Modern Data Analytics Solutions on AWS course cover?
This four-day bootcamp covers the complete AWS data analytics lifecycle: building data lakes with AWS Glue and Lake Formation, batch processing on Amazon EMR with Apache Spark and Hive, data warehousing with Amazon Redshift, and real-time streaming with Amazon Kinesis and Amazon MSK. IT Dojo delivers it as live instructor-led training with an emphasis on practical skills for government and DoD professionals.
How long is IT Dojo's Building Modern Data Analytics Solutions on AWS training?
IT Dojo's Building Modern Data Analytics Solutions on AWS training is 4 Days. It is available as live remote online instruction or on-site at your facility. All sessions are instructor-led with small class sizes to ensure individual attention.
Is this course available as live remote online training?
Yes. IT Dojo offers Building Modern Data Analytics Solutions on AWS as live remote online training. A certified instructor leads the session in real time. Students interact via chat or microphone. Classes are kept small (typically no more than 16 students) to ensure engagement. On-site delivery at your government facility or contractor location is also available.
What prerequisites are recommended before this course?
We recommend that attendees of this course have: AWS Technical Essentials Architecting on AWS.
Does IT Dojo offer this training on-site at government or DoD facilities?
Yes. IT Dojo delivers Building Modern Data Analytics Solutions on AWS on-site at government agencies, DoD commands, military installations, and contractor facilities. On-site training is ideal for teams of four or more and can be customized to your organization's specific environment and mission requirements. Contact IT Dojo to schedule.
How do I register for this course?
IT Dojo training is employer sponsored. Your organization registers and pays for seats. To schedule Building Modern Data Analytics Solutions on AWS for your team, contact IT Dojo via the Request Training form or call 757-216-3656. IT Dojo will work with your contracting officer, training coordinator, or program office to set up the course.