BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260422T143139Z
LOCATION:B309
DTSTART;TZID=America/New_York:20241119T133000
DTEND;TZID=America/New_York:20241119T140000
UID:submissions.supercomputing.org_SC24_sess403_pap440@linklings.com
SUMMARY:Automated Code Generation of High-Order Stencils for a Dataflow Ar
 chitecture
DESCRIPTION:Ryuichi Sai, John Mellor-Crummey, and Jinfan Xu (Rice Universi
 ty) and Mauricio Araya-Polo (TotalEnergies E&P Research and Technology USA
 , LLC)\n\nFinite-difference methods based on high-order stencils are widel
 y used in seismic simulations, weather forecasting, and computational flui
 d dynamics. Recently, multiple research groups have begun exploring the us
 e of dataflow architectures, such as Cerebras' wafer-scale engine, to acce
 lerate stencil computations. However, implementations of stencil computati
 ons for dataflow architectures must address unique challenges, such as man
 aging the routing of data communications and accommodating a significantly
  constrained memory footprint. These make hand-crafting code for a dataflo
 w architecture difficult and time-consuming. This paper describes a framew
 ork for developing portable, high-performance implementations of stencil c
 omputations for modern node architectures. The paper focuses on code gener
 ation strategies for the Cerebras wafer-scale engine, including code gener
 ation of router configurations and sequencing of communication for high-or
 der stencils. A 25-point star-shaped stencil written using our tool is 7x 
 shorter than hand-crafted code written in Cerebras Software Language (CSL)
 , and it delivers comparable performance to manually written code.\n\nTag:
  Accelerators, Compilers, Embedded and/or Reconfigurable Systems, Linear A
 lgebra, Performance Evaluation and/or Optimization Tools\n\nRegistration C
 ategory: Tech Program Reg Pass\n\nSession Chair: Keita Teranishi (Oak Ridg
 e National Laboratory (ORNL))\n\n
END:VEVENT
END:VCALENDAR
