Establishing a Central Data Function at a Public Agency
NYC Open Data Week – March 26, 2026
VIDEO | AUDIO | RECAP EN / ES / FR | INFO | INDEX
Speakers: Andrew Kuziemko - Deputy Chief, Data & Analytics, MTA
Moderator: Lisa Mae Fiedler - Manager, Open Data, MTA
At NYC Open Data Week 2026, Andrew Kuziemko outlined how the Metropolitan Transportation Authority (MTA) built a centralized Data & Analytics function to modernize analytics, reduce silos, and create shared infrastructure for data-driven decision-making across the agency. The session focused less on technical theory and more on organizational change, practical governance, and the realities of building data systems inside a large public agency.
Why Create a Central Data Function?
Kuziemko explained that the MTA historically operated with highly fragmented data systems. Different departments maintained their own databases and reporting processes, often with little interoperability. Analysts seeking information had to negotiate access with database owners or IT staff, making even routine analysis slow and frustrating.
He argued that a centralized data team can solve several recurring organizational problems:
Maintain a shared data repository for analysts
Create a “single source of truth” for metrics
Standardize analytical practices
Promote transparency and data sharing
Improve productivity through automation
Provide common analytics tools and tutorials
Reduce manual reporting pipelines
Support modern governance practices
Rather than every department reinventing its own reporting process, the goal was to establish reusable infrastructure that analysts throughout the organization could rely on.
Managing From Data
Kuziemko framed analytics as part of a broader operational feedback loop:
Operational systems generate data
Analysts extract insights
Managers make decisions
Operational changes occur
Results appear back in the data
He stressed that analytics itself should be treated as a professional discipline, comparable to engineering or operations management. Analysts must acquire data, clean and transform it, visualize findings, and communicate results to decision-makers. Much of this pipeline, he noted, can and should be automated.
The central data team’s role is to eliminate friction between analysts and operational databases. Instead of every analyst becoming an “amateur data engineer,” the centralized group handles extraction, transformation, documentation, and infrastructure, making clean data readily available through a common platform.
Formation of the MTA Data & Analytics Team
Kuziemko described how the initiative emerged from a crisis in legacy reporting infrastructure. Around 2021, he inherited responsibility for major MTA performance reporting systems after a retirement left critical reporting pipelines unstable and poorly documented.
Many systems relied on outdated practices:
On-premises virtual machines
Cron jobs
Manual interventions
Poor security practices
Infrastructure hidden “under someone’s desk”
Analysts manually restarting broken processes
Using these weaknesses as justification, Kuziemko secured modest additional staffing and support to rebuild the reporting environment on cloud infrastructure. The team launched officially in January 2022 as “MTA Data and Analytics,” positioned centrally within MTA headquarters rather than inside a single operating agency.
The Open Data program was later folded into the team’s responsibilities, including management of public-facing datasets and performance reporting.
Data Lake Architecture and Shared Infrastructure
The MTA team built a centralized data lake in Azure to consolidate operational data and analytical pipelines. Thousands of tables are now maintained in the environment, organized into a broadly accessible “core” zone and more restricted topic-specific zones for sensitive information such as HR or legal data.
The core zone contains non-sensitive data broadly accessible to MTA analysts, while restricted zones manage more sensitive datasets. Kuziemko emphasized that much analytically useful public-sector data does not contain personally identifiable information, making broad sharing feasible.
The team deliberately avoided building overly complicated governance structures before proving the concept. Instead, they prioritized practical delivery using non-sensitive datasets, allowing trust and adoption to develop gradually.
Defining New Technical Roles
One of the most important lessons, according to Kuziemko, was distinguishing analysts from data engineers. Before the centralized effort, analysts routinely built fragile “shadow infrastructure” using Excel macros, Power BI workarounds, or ad hoc Python scripts to automate repetitive work.
The team formalized several new roles:
Data Engineers
Analytics Engineers
Data Scientists
Data engineers maintain infrastructure and pipelines. Analytics engineers handle extraction and cleaning workflows. Data scientists focus on business questions, modeling, visualization, and communication with operational teams.
This specialization allowed analysts to focus more on interpretation and decision support instead of pipeline maintenance.
Strategic Early Projects
Kuziemko advised organizations to select initial projects carefully. Ideal early projects should:
Be highly visible
Use large datasets
Solve recognizable organizational problems
Avoid highly sensitive information
Demonstrate quick value
At the MTA, initial projects included ridership statistics, subway on-time performance, and congestion pricing reporting. Because these metrics were already operationally critical and publicly scrutinized, improvements were immediately visible across the organization.
The congestion pricing rollout became a particularly important example. The team rapidly built automated pipelines to publish daily traffic and transit data shortly after the program launched, helping defend the initiative publicly with empirical evidence about reduced traffic and improved bus speeds.
Working With IT Instead of Against It
Although many departments historically bypassed IT in order to move faster, Kuziemko strongly encouraged collaboration with IT departments rather than isolation. The MTA team partnered with IT to establish Azure infrastructure, clear open-source tools, and ensure transparency in platform operations.
He acknowledged that relationships with IT can initially feel frustrating, especially where legacy systems dominate, but argued that sustainable infrastructure requires long-term partnership rather than parallel shadow systems.
Executive Buy-In and Organizational Change
Kuziemko stressed that executive support depends on translating technical frustrations into business frustrations executives recognize. Executives may not personally struggle with data access, but they do experience:
Slow answers
Conflicting reports
Inconsistent metrics
Delayed decision-making
The central data function was presented not as a technology upgrade, but as a way to improve responsiveness, consistency, and organizational efficiency.
Importantly, the investment required was relatively modest. The MTA initially added only about five data engineering staff positions and began with a comparatively inexpensive cloud budget. Most other roles were converted from existing analyst positions.
Cultural Change and Early Adopters
A recurring theme throughout the session was that organizational change happens more effectively through demonstrated usefulness than through mandates. Kuziemko described how formal executive memos requiring cooperation had limited effect, while successful early projects and enthusiastic “early adopters” gradually created bottom-up demand for the new infrastructure.
Instead of forcing resistant departments to participate immediately, the team focused on cooperative partners and visible successes. Over time, more groups requested inclusion after seeing the benefits others received.
Open Data and Transparency
The Open Data program played a major role in reinforcing the culture of transparency. Public datasets and dashboards demonstrated the practical value of centralized infrastructure while also encouraging internal sharing norms.
Lisa Mae Fiedler noted that many non-technical users simply access MTA information through the public Open Data portal rather than directly querying the internal data lake. The team is also exploring tools like CKAN to make smaller datasets more accessible to less technical staff.
Conclusion
The session presented the MTA’s data modernization effort as an exercise in organizational design as much as technical engineering. Kuziemko emphasized that successful data transformation depends on creating shared infrastructure, defining specialized roles, building trust incrementally, and demonstrating immediate operational value.
Rather than beginning with ambitious governance frameworks or highly sensitive data, the MTA focused on solving visible operational problems with scalable infrastructure and practical automation. Over four years, this approach evolved into a centralized data ecosystem supporting analytics, Open Data publication, performance reporting, and agency-wide decision-making.
RESOURCES
Establishing a Central Data Function at a Public Agency — the NYC Open Data Week event page for this session
Lessons Learned in Starting a Central Data Team — Andy Kuziemko’s written companion to this talk, covering the same eight lessons
Lessons Learned in Managing the MTA’s Open Data Program — Lisa Mae Fiedler’s guide to launching a government open data program
How We Build Analytics at Scale at the MTA — Mike Kutzma’s deep dive on the data lake, Airflow orchestration, and pipeline architecture
MTA Open Data Program — the program Lisa Mae Fiedler runs, with annual plans and team contact
MTA Metrics Site — the public dashboard built on open data, including subway on-time performance and ridership
MTA Data on the NYS Open Data Portal — the 230+ open datasets the data lake feeds via Reverse ETL
MTA Bus Automated Camera Enforcement Violations — the ABLE/ACE dataset Lisa worked on, referenced in the talk
Automated Camera Enforcement (ACE) — background on the bus-mounted camera program discussed
Go ACE! (Streetsblog) — the kind of press coverage built on MTA open data that Andy cited


