Platform Overlays: Enabling In-Network Stream Processing
Download
Report
Transcript Platform Overlays: Enabling In-Network Stream Processing
Green Clouds – Power Consumption
as a First Order Criterion
Karsten Schwan, Sudhakar Yalamanchili, Ada Gavrilovska,
Hrishikesh Amur, Bhavani Krishnan, Surabhi Diwan,
Nikhil Sathe, Minki Lee, Saibal Mukopadhyay, …
CERCS
Yogendra Joshi, Pramod Kumar, Emad Samadiani
CEETHERM
Georgia Institute of Technology
http://img.all2all.net/main.php?g2_itemId=157
• An eco-system of
projects
addressing
multiple stack
layers
• Multiple faculty
and students
involved
Datacenter and beyond:
design, IT management,
HVAC control… (ME, SCS,
OIT…)
Rack: mechanical design,
thermal and airflow analysis,
VPTokens, OS and management
(ME, SCS)
Board: VirtualPower,
scheduling/scaling/operating
system… (SCS, ME, ECE)
• ECE, ME, CS
Chip and Package: power
multiplexing, spatiotemporal
migration (SCS, ECE)
Circuit level: DVFS, power
states, clock gating (ECE)
Power distribution and delivery (ECE)
Green Computing
Initiative
Modeling and Control Across
Entire Stack
•
Seek a fundamental understanding of relationships between
performance, power distribution, energy consumption, heat
generation and cooling technologies at all levels of the stack
•
Develop, model, and assess (new) principles for energy and
thermal management
– Coordinated management across the entire stack
– Example concepts
• Couple cooling and workload generation: thermal flow
control to respond to load conditions
• Couple power distribution and workload generation:
adapt to power capacity (time of day?)
• Pro-active spatio-temporal migration driven by physics of
heat generation/flow rather than reactive sensor driven
techniques
Sample Projects
• Understanding of power distribution
opportunities at the chip level
• Platform-level coordination of power
management methods
– DVFS, scheduling, idle states
– Understanding of impact of these approaches
• Distributed power management methods
• IT & environmental factor management
– Temperature and air velocity, HVAC control
• Management architecture for virutalized
platforms
Thermal and Power Scaling
Limits-On-Chip
Temperature Limited
Performance
Power Limited Performance
Mukhopadhyay and Yalamanchili (2008)
Unmanaged Thermal Behavior
Managed Thermal Behavior
(multiplexed power)
Spatial gradient:
10.5°@0.75mm
Temporal
gradient:
2.5°@100Kcyle
s gradient:
Spatial
75°
70°
45°
2.5°@0.75mm
Temporal gradient:
1.99°@100Kcyle
s
75°
70°
• 64 on-tiles
• 256 total tiles
• 100K time slice
interval@3GHz
45°
30 ms
Courtesy: Nikil Sathe
30 ms
The Need for Feedback
Thermal profile
Co-exploration of thermal
management/architecture
management
Spatiotemporal
migration
core
Local
CacheMemory
core
core
Local
Local
CacheMemory CacheMemory
core
Local
CacheMemory
Co-design power
distribution/architecture
management@chip and
multi-chip
Power distribution
network
Towards Integrated Platform Power
Management
• Multiple techniques exist for power
management on platforms.
VirtualPower: Coordinated Power
Management
Dom0
VPM Mechanisms
VPM
Rules
PM
Policy
Application
VPM States OS
Application
VPM States OS
VPM Channel
Hypervisor
Hardware
• Coordinated power management
(DVFS) + load management
(migration) + CPU management
(credit based soft scaling) ->
cumulative reduction of 34% in
power resources without SLA
degradation of RUBis
benchmark
•
PM
Policy
(R. Nathuji, K. Schwan, SOSP0)
Platform-level Power Management
Methods: Costs and Opportunities
• Goals
– identify and quantify the
reasons for performance
degradation associated with
each method for power
management
– To be able to estimate the
power savings from each
methods
• Use this information to
support better runtime
management decisions
• Current focus: DVFS
– Bounds on degradation based
on system profiling and
runtime performance couter
information
• Other methods next,
including idle states
Virtualization Support for Power
Budgeting
VM4
VM1
VM3
VM5
VM2
Goal: Develop
system support for
improving
Power
Power
aggregate QoS/performance of VMs in distributed
power budgeted environments
Platform power cap
directly affects VM
performance impact
Group level power
budget distributed
amongst underlying
platforms based on
system utilization, static
priority, etc
Power
Power
VPMTokens Benefits
Norm alized Ap p licat ion QoS
1
VPM input allows system to
dynamically move budget
from low rate to high rate
VM, reducing overall
performance impact within
budget constraints
0.9
0.8
0.7
0.6
0.5
0.4
100%
Trans (low rate)
Trans (high rate)
Trans (high rate
w/VPM channel)
Without VPM channel
feedback high rate
transaction application
experiences QoS impact
90%
80%
70%
60%
Plat form Bu d g et (%)
• Experiment: High rate and low rate transaction VMs running on P4 platform
• Both VMs have same utility value, therefore equal allocation without VPM
channel input
• (R. Nathuji, K. Schwan, HPDC08)
CoolIT: Coordinated IT and
environmental/thermal management
Use of power budgeting in IT
• Efficient operating point for
cooling infrastructure implies
load constraints
• Power capping capabilities
enable dynamic compliance
when DPM alone cannot meet
constraints
Energy tradeoff
• Increased airflow rates provide improved ability to
dissipate server heat loads
• Airflow required to meet maximum inlet
temperature varies based upon server loads and
physical distribution of heat
C
R
A
C
R1
R2
R3
R4
Cold Aisle
R5
R6
R7
R8
CoolIT Approach
• Objectives:
– VM load distribution strategy:
• Inlet temperature in the data center should be below
32C (or target temperature)
• Server power consumption and cooling power is limited
• Model output:
300
250
Power (in watts)
– Heat load within limits based on cooling capacity
– Hot-spot avoidance in the face of load and platform
heterogeneity
200
150
– target utilization for different platforms based on inlet
100
temperature, air velocity, average utilization
• Offline profiles:
– power vs CPU utilization for different architectures
– other resources next
• Dynamic load management:
– Currently first fit bin-packing algorithm
– Driven by management server
• Input from mgt domains and thermocouples sensors
– OSIsoft PI server
• dom0 -> PI server SNMP-based infrastructure
• can also drive VM migration
50
0
CPU Utilization
Management Architecture
• Management brokers
– make and enforce ‘localized’ management decisions
• within VMs
• VMM-level – CPU scheduling, allocation of memory or device
resources, ..
• at hardware level
• Management channels
– enable inter-broker
coordination through
well-defined interfaces
– event and shared
memory based
• Management VMs
– platform wide policies
and cross-platform
coordination
CERCS Distributed Cloud
Infrastructure