Data Science and the Knowledge Discovery AdventureThis talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)
Document ID
20220010003
Acquisition Source
Ames Research Center
Document Type
Presentation
Authors
Bryan Matthews (Wyle (United States) El Segundo, California, United States)
Date Acquired
June 28, 2022
Subject Category
Computer SystemsStatistics And Probability
Meeting Information
Meeting: University of California Riverside Summer Fellowship Program Talk