Level 1 Calorimeter Trigger - University of Wisconsin
Download
Report
Transcript Level 1 Calorimeter Trigger - University of Wisconsin
Grid Computing - A Primer
Sridhara Dasu, Department of Physics, U. Wisconsin
• Grid Computing
– What is the buzz all about?
– What is the promise?
• My Perspective
– What is in it for me?
– How is it working for us?
• In UW-Madison
• And, beyond …
Acknowledgements:
Condor Team
GLOBUS Team
I.Foster/Argonne
M.Livny/Wisconsin
D.Bradley/Wisconsin
• Conclusion
– Why should you be interested?
– What are the consequences for you?
July 17, 2015
Sridhara Dasu
1
Grid Computing is
in the News …
July 17, 2015
Sridhara Dasu
2
Grid Projects Are Ubiquitous
The Opportunity (or Challenge):
Computational Cornucopia
• Abundant computation, data, bandwidth
– In many fields, too much data—not too little
– Simulations of unprecedented accuracy
– Ubiquitous internet distance not a barrier
• But as a consequence
– Rate of change accelerates
– Complex problems multidisciplinary distributed
teams & sharing of resources & expertise
– Without infrastructure, you can’t compete
July 17, 2015
Sridhara Dasu
4
Why Distributed Teams
Are Important
• Increasingly challenging & complex problems
– Particle physics, Global change, Cosmology, Life
sciences
– Manufacturing, Mineral exploration
– Film production, Game development, …
• Required expertise & resources also distributed
–
–
–
–
July 17, 2015
People
Computational capability
Data
Sensors
Sridhara Dasu
5
The Grid
“Resource sharing & coordinated problem
solving in dynamic … virtual
http://www.mkp.com/mk/default.asp?isbn=1558609334
organizations”
1. Enable integration of distributed service & resources
2. Using general-purpose protocols & infrastructure
3. To achieve useful qualities of service
“The Anatomy of the Grid”,
Foster, Kesselman, Tuecke, 2001
Sridhara Dasu
July 17, 2015
6
What is a Grid?
• The key criteria:
– Coordinated distributed resources …
– Uses standard, open, general-purpose protocols
and interfaces …
– Deliver non-trivial qualities of service.
• What is not a Grid?
– A cluster, a network attached storage device, a
scientific instrument, a network, etc.
– Each is an important component of a Grid, but by itself
does not constitute a Grid
July 17, 2015
Sridhara Dasu
7
Why Should You Care?
1) Grid is a promising technology [Vision]
– It ushers in a virtualized, collaborative, distributed
world
2) Grids are being commissioned now [Reality]
– Grids are built (not bought), but are delivering real
benefits in academic and commercial settings
3) An open Grid is to your advantage [Future]
– Standards are being defined now that will determine
the future of this technology
July 17, 2015
Sridhara Dasu
8
Quality, economies of scale
The Power Grid:
On-Demand Access to Electricity
Decouple production &
consumption, enabling
• On-demand access
• Economies of scale
• Consumer flexibility
• New devices
Time
July 17, 2015
Sridhara Dasu
9
But Computing Isn’t
Really Like Electricity!
• How about “access computing resources like
we access Web content”?
– We have no idea where a website is, or on what
computer or operating system it runs
Two interrelated opportunities
1) Enhance economy, flexibility, access by
virtualizing computing resources
2) Deliver entirely new capabilities by
integrating distributed resources
July 17, 2015
Sridhara Dasu
10
Virtualization
• Automatically connect
applications to services
• Dynamic & intelligent
provisioning
Applications:
Delivery
Application Virtualization
Application
Services:
Distribution
Infrastructure Virtualization
Servers:
Execution
• Dynamic & intelligent
provisioning
• Automatic failover
Source:
The Grid: Blueprint for a New
Computing
Infrastructure (2nd Edition), 2004
July 17, 2015
Sridhara
Dasu
11
Local Clusters to Global Grids
Cluster Grid
July 17, 2015
Enterprise Grid
Sridhara Dasu
Global Grid
12
Mission Criticality
Grid Deployment Trends
Corporate
Scientific
Department
July 17, 2015
Enterprise
Collaboration
Sridhara Dasu
Internet
13
Transparent Service
Grid
Utility
Computing
Autonomic
Computing
ServiceOriented
Architecture
Webster says: Autonomic = acting or occurring involuntarily
<autonomic reflexes>
July 17, 2015
Sridhara Dasu
14
Layers of Grid Architecture
July 17, 2015
Sridhara Dasu
15
Multidisciplinary Teams:
Problem Solving in the 21st Century
• Teams organized around common goals
– Communities: “Virtual organizations”
• With diverse membership & capabilities
– Heterogeneity is a strength not a weakness
• And geographic and political distribution
– No location/organization possesses all required
skills and resources
• Must adapt as a function of the situation
– Adjust membership, reallocate responsibilities,
renegotiate resources
July 17, 2015
Sridhara Dasu
16
Challenging Technical
Requirements
• Dynamic formation and management of virtual
organizations
• Discovery & online negotiation of access to services:
who, what, why, when, how
• Configuration of applications and systems able to
deliver multiple qualities of service
• Autonomic management of distributed infrastructures,
services, and applications
• Management of distributed state
• Open, extensible, evolvable infrastructure
July 17, 2015
Sridhara Dasu
17
The Globus Project™
Making Grid computing a reality (since 1996)
• Close collaboration with real Grid projects in science
and industry
• The Globus Toolkit®: Open source software base for
building Grid infrastructure and applications
• Development and promotion of standard Grid
protocols to enable interoperability and shared
infrastructure
• Development and promotion of standard Grid software
APIs to enable portability and code sharing
• Global Grid Forum: We co-founded GGF to foster Grid
standardization and community
July 17, 2015
Sridhara Dasu
18
Globus Toolkit 2
Key Protocols
• The Globus Toolkit v2 (GT2)
centers around four key protocols
– Connectivity layer:
• Security: Grid Security Infrastructure (GSI)
– Resource layer:
• Resource Management: Grid Resource Allocation
Management (GRAM)
• Information Services: Grid Resource Information Protocol
(GRIP)
• Data Transfer: Grid File Transfer Protocol (GridFTP)
• Also key collective layer protocols
– Info Services, Replica Management, etc.
July 17, 2015
Sridhara Dasu
19
Resource Management
UW Condor Project - Miron Livny’s group
(http://www.cs.wisc.edu/condor)
– Predates Globus
– High throughput computing on commodity resources
– Successful enterprise level deployment
•
•
•
•
•
•
•
July 17, 2015
UW Computer Science Condor pool
UW Condor pools in other departments
INFN/Italy pools
Inter-pool flocking
…
Also, some industrial users
…
Sridhara Dasu
20
The Layers of Condor
Complete solution for resource management
Application
Application Agent
Submit
(client)
Customer Agent
Matchmaker
Owner Agent
Remote Execution Agent
Execute
(service)
Local Resource Manager
Resource
July 17, 2015
Sridhara Dasu
21
A Grid Job
• Must be able to run in the background: no
interactive input, windows, GUI, etc.
• Can still use STDIN, STDOUT, and
STDERR (the keyboard and the screen),
but files are used for these instead of the
actual devices
• Organize data files, input/output
July 17, 2015
Sridhara Dasu
22
Condor Universes
• The Standard Universe
– Check-points executable state
– Job migration to other resources to continue execution
– Transparent IO redirection to user submit machines
– Robust against resource preemption for higher priority tasks +
resource failures
– Limitations on applications (e.g., shlib, MT)
• The Vanilla Universe
– Traditional batch jobs with no limitations
– External solutions for IO redirection
– Not robust against preemption or resource failures
• The Globus Universe (new)
– Adapted to emerging Grid standards
– Part of Globus Toolkit
July 17, 2015
Sridhara Dasu
23
Condor-G: Globus + Condor
Globus
•
•
•
Condor
middleware deployed across
entire Grid
remote access to computational
resources
dependable, robust data
transfer
July 17, 2015
•
•
•
Sridhara Dasu
job scheduling across multiple
resources
strong fault tolerance with
checkpointing and migration
layered over Globus as
“personal batch system” for the
Grid
24
Condor-G
User/Application
Condor
Grid
Globus Toolkit
Condor
Fabric (processing, storage, communication)
July 17, 2015
Sridhara Dasu
25
Creating a Submit Description File
• A plain ASCII text file
• Tells Condor-G about your job:
– Which executable, grid site, input, output and
error files to use, command-line arguments,
environment variables, etc.
• Can describe many jobs at once (a
“cluster”) each with different input,
arguments, output, etc.
July 17, 2015
Sridhara Dasu
26
Simple Submit Description File
# Simple condor_submit input file
# (Lines beginning with # are comments)
# NOTE: the words on the left side are not
#
case sensitive, but filenames are!
Universe
= globus
GlobusScheduler = host.domain.edu/jobmanager
Executable = my_job
Queue
July 17, 2015
Sridhara Dasu
27
Running condor_submit
• You give condor_submit the name of the submit file
you have created
• condor_submit parses the file, checks for errors, and
creates a “ClassAd” that describes your job(s)
• Sends your job’s ClassAd(s) and executable to the
Condor-G schedd, which stores the job in its queue
– Atomic operation, two-phase commit
• View the queue with condor_q
July 17, 2015
Sridhara Dasu
28
condor_submit sequence
Globus Resource
Condor_q
Gate Keeper
Condor_submit
Condor-G
Local Job
Scheduler
Condor-G
July 17, 2015
Sridhara Dasu
29
Running condor_submit
% condor_submit my_job.submit-file
Submitting job(s).
1 job(s) submitted to cluster 1.
% condor_q
-- Submitter: perdita.cs.wisc.edu : <128.105.165.34:1027> :
ID
OWNER
SUBMITTED
RUN_TIME ST PRI SIZE CMD
1.0
frieda
6/16 06:52
0+00:00:00 I 0
0.0 my_job
1 jobs; 1 idle, 0 running, 0 held
%
July 17, 2015
Sridhara Dasu
30
DAGMan
• Directed Acyclic Graph Manager
• DAGMan allows you to specify the
dependencies between your Condor-G jobs,
so it can manage them automatically for you.
• (e.g., “Don’t run job “B” until job “A” has
completed successfully.”)
July 17, 2015
Sridhara Dasu
31
What is a DAG?
• A DAG is the data structure
used by DAGMan to represent
these dependencies.
• Each job is a “node” in the
DAG.
Job A
Job B
• Each node can have any
number of “parent” or “children”
nodes – as long as there are no
loops!
July 17, 2015
Sridhara Dasu
Job C
Job D
32
Defining a DAG
• A DAG is defined by a .dag file, listing each of its
nodes and their dependencies:
Job A
# diamond.dag
Job A a.sub
Job B b.sub
Job C c.sub
Job D d.sub
Parent A Child B C
Parent B C Child D
Job B
Job C
Job D
• each node will run the Condor-G job specified by its
accompanying Condor submit file
July 17, 2015
Sridhara Dasu
33
What about Data?
Data Placement* (DaP) must be an integral part
of the end-to-end solution
Stork (Another UW-Computer Science Product)
– Schedules, runs, monitors, and manages Data
Placement (DaP) jobs in a heterogeneous Grid
environment & ensures that they complete.
– What Condor (G) means for computational
jobs, Stork means the same for DaP jobs.
– Just submit a bunch of DaP jobs and then relax..
– Interoperates with various storage services
* Space management and Data transfer
July 17, 2015
Sridhara Dasu
34
Full Condor-G Capabilities
Planner(s)
DAGMan
Condor-G
Stork
(compute)
Gate
Keeper
StartD
SRB
SRM
July 17, 2015
(DaP)
NeST GridFTP
Sridhara Dasu
RFT
35
UW “Enterprise Level” Grid
• Condor pool at CS
– 1000 ~1GHz Intel CPUs
• Condor pools at various departments
– 100 ~2.4 GHz Intel CPUs at Physics, etc.
– New: Grid Laboratory of Wisconsin
• Condor jobs flock from various departments
to CS Pool as needed
• Excellent utilization
– Especially when the Condor Standard Universe is
used
• Premption, Checkpointing, Job Migration
July 17, 2015
Sridhara Dasu
36
Grid Laboratory of Wisconsin
2003 Initiative funded by NSF/UW
Six GLOW Sites
• Computational Genomics, Chemistry
• Amanda, Ice-cube, Physics/Space Science
• High Energy Physics/CMS, Physics
• Materials by Design, Chemical Engineering
• Radiation Therapy, Medical Physics
• Computer Science
Phase-1 already has ~300 Xeon CPUs
Expect to grow to about 700 CPUs + 100 TB disk
July 17, 2015
Sridhara Dasu
37
Condor/GLOW Ideas
• Exploit commodity hardware for high
throughput computing
– The base hardware is the same at all sites
– Local configuration optimization as needed
• e.g., Number of CPU elements vs storage elements
– Must meet global requirements
• It turns out that our initial assessment calls for almost
identical configuration at all sites
• Managed locally at 6 sites
– Shared globally across all sites
– Higher priority for local jobs
July 17, 2015
Sridhara Dasu
38
The Large Hadron Collider
July 17, 2015
Sridhara Dasu
39
The Large Hadron Collider
Building and commissioning the
accelerator and detectors, and extracting
interesting physics out of this massive
data sample is a big challenge.
July 17, 2015
Sridhara Dasu
40
Event Filtering Before Archival
Output:
1MB/event @100 Hz
Petabyte per year
July 17, 2015
Sridhara Dasu
41
Analysis Teams + Resources
Input:
~109 events (petabyte databases)
Complex
algorithms
developed by
collaborating
physicists
Output:
Publications with ~100s of selected events
July 17, 2015
Sridhara Dasu
42
Simulation: Early Grid Deployment
•
•
•
Detailed simulations necessary
– Large numbers of background events need to be simulated
• Dominated by fluctuations of tails
Computation scale
– Background events occur on every crossing - 40 MHz
• Up to 10 minutes on a 1 GHz CPU to simulate full event
• 2 x 109 s CPU time to simulate 1 s of LHC operation
• Requires 1000 CPUs running for 1 month
– CMS has large number of detector channels, 108
• Each event requires 1-10 MB storage space
• 32-320 TB needed for 1 s of LHC operation
– Optimizing CPU and data storage
• Simulate in bins and reuse some data
Pleasantly parallel application
– Ideal Grid testbed candidate
• Used UW “enterprise level” classic Condor grid successfully
• With Grid2003 used nation wide Globus/Condor-G based true grid
July 17, 2015
Sridhara Dasu
43
Tapping UW “Enterprise Level” Grid
2000
1500
2001
2002
1000
2003
500
We tapped resources on the UW
campus opportunistically
2004
0
# of 1-GH z Intel CPU s
10000000
8000000
6000000
4000000
2000000
0
# of events simulated
We produced more events in
2003 than most other CMS
collaborators - because of
using our UW enterprise
level grid and condor
standard universe!
2004 numbers are through March, and were also running our new C++ simulation code that is a factor of 2 slower.
We have typically used less than 50% of available resources and ran for about 30% of the year.
July 17, 2015
Sridhara Dasu
44
Tapping Global Grid : Grid3
July 17, 2015
Sridhara Dasu
45
Cost Savings from Grids
• The size of cost savings from
grids will come in two waves:
– First from the adoption of
clusters
– Then from the adoption of
Enterprise Grids
• Firms using Clusters estimate
that cost savings will be small at
first, but will grow to 15% to 30%
savings in IT Costs in 2005-2008.
• Firms planning to use Enterprise
Grids estimate that they will
experience a second wave of
benefits. Savings will grow to
15% to 30% by 2007-2010.
July 17, 2015
Source: Robert Cohen, “Grid Computing: Projected
Impact on North Carolina’s Economy & Broadband
Use through 2010,” Rural Internet Access
Authority, September 2003. http://www.e-nc.org
Sridhara Dasu
46
Grid drawbacks being
addressed now
• Low utilization of enterprise resources
• High cost of provisioning for peak
demand
• Inadequate resources prevent use of
advanced applications
• Lack of information integration
July 17, 2015
Sridhara Dasu
47
Cyberinfrastructure & VOs
Relevance Far Beyond Science
1) Virtualization of information technology
– From vertical silos to on-demand access
– Improve efficiency of delivery, increase flexibility of
use
– E.g., financial services, e-commerce
2) New applications, products, & services
enabled by much computation & data
– Media, life sciences, manufacturing, seismic
exploration, online gaming, etc., etc., etc.
July 17, 2015
Sridhara Dasu
48
The Value of Grid Computing:
IBM Perspective
Increased
Efficiency
Higher Quality
of Service
Increased
Productivity
& ROI
Reduced
Complexity
& Cost
Improved
Resiliency
July 17, 2015
Sridhara Dasu
49
Grids: HP Perspective
virtual data center
computing
utility or
GRID
value
programmable data center
switch
compute fabric storage
grid-enabled
systems
clusters
UDC
Tru64, HP-UX,
Linux
Open VMS clusters,
TruCluster, MC
ServiceGuard
July 17, 2015
today
shared, traded resources
Sridhara Dasu
50
Grid Vision, Marketing, and Reality
• Vision
– Computing resources can be shared like content on
the Web
• Marketing
– Have we got a Grid for you!
• [Data, compute, knowledge, information, desktop, PC,
enterprise, cluster, …]
• Reality
– Commercial products mostly non-interoperable
– Open source tools offer de facto standards, but are
also far from a complete solution
July 17, 2015
Sridhara Dasu
51
Standards Matter!
• Open, standard protocols
–
–
–
–
Enable interoperability
Avoid product/vendor lock-in
Enable innovation/competition on end points
Enable ubiquity
• In Grid space, must address how we
– Describe, discover, & access resources
– Monitor, manage, & coordinate, resources
– Account & charge for resources
For many different types of resource
July 17, 2015
Sridhara Dasu
52
Increased functionality,
standardization
Developing Grid Standards
Managed shared
virtual systems
Research
Open Grid
Services Arch
Web services, etc.
Internet
standards
Custom
solutions
1990
July 17, 2015
Real standards
Multiple implementations
Globus Toolkit
Defacto standard
Single implementation
1995
2000
Sridhara Dasu
2005
2010
53
Open Grid Services Architecture
Adopt service-oriented architecture
– Key to virtualization, discovery, composition, localremote transparency
+ Standard service description & access
– Leverage industry standard Web services
+ Distributed service management protocols
– A “component model for Web services”
= A framework for creating, managing, &
delivering interoperable services
“The Physiology of the Grid: An Open Grid Services Architecture for
Distributed
Systems Integration”,
Foster,
July 17, 2015
Sridhara
Dasu Kesselman, Nick, Tuecke, 2002
54
Grid & Web Services
• Grid and Web Services are merging
– Grid is an aggressive use case of Web Services
• Web Services standards landscape is in flux
– OGSI/A will need to evolve with it
– Uncertain status of security & policy standards
continues to be a big source of concern
• Grid services standards landscape heating up
• W3C, OASIS, GGF are key standards orgs
• Open source software important for adoption
July 17, 2015
Sridhara Dasu
55
OGSA Status: Implementations
• Globus Toolkit v3: Linux for the Grid
– Open source middleware, commercial support
– A range of computation & data management,
registry, and security functions
• Some nice announced OGSI-based products
– IBM, Avaki, Platform, Sun, NEC, HP, UD, Entropia,
DataSynapse, Insors, Oracle, etc.
– Read the fine print: “Intent is to use OGSI-based
products,” “OGSA-compliant software,” “Embraces
fundamental OGSA concepts”
July 17, 2015
Sridhara Dasu
56
Why Should You Care?
1) Grid is a promising technology [Vision]
– It ushers in a virtualized, collaborative, distributed
world
2) Grids are being commissioned now [Reality]
– Grids are built (not bought), but are delivering real
benefits in academic and commercial settings
3) An open Grid is to your advantage [Future]
– Standards are being defined now that will determine
the future of this technology
July 17, 2015
Sridhara Dasu
57
Consequences for Network
• Increased bandwidth use
– Data transport for distributed collaborative
applications are likely to be larger than web
browsing
• Demands on Quality of Service
– As Grid computing becomes integral part of user
needs QoS requirements go up
• Access, Authentication, Auditing
– Issues are being addressed adequately as Grids
are being commissioned
• However, there are always malicious hackers
• And, there is also naïve users unintentionally causing
lockup by demanding excessive resources on the grid
July 17, 2015
Sridhara Dasu
58
Pop Quiz: The Grid Is …
a) A collaboration & resource sharing
infrastructure for scientific applications
b) A distributed service integration and
management technology
c) A disruptive technology that enables a
virtualized, collaborative, distributed world
d) An open source technology & community
e) A marketing slogan
f) All of the above
July 17, 2015
Sridhara Dasu
59
To Learn More
Deep tech & strategic info
Jan 20-23, 2004, San Fran
www.globusworld.org
Working, research groups
Three meetings a year
www.ggf.org
Weekly email & web
newsletter
www.gridtoday.com
2nd Edition
www.mkp.com/grid2
Commercial conf & expo
May 24-26, Philadelphia
www.gridtoday.com
July 17, 2015
Sridhara Dasu
60