大数据算法(专业核心)
学分:3.0
本解析基于笔者在丁虎班-2026春课程体验所得,对于不同时间、不同授课教师,课程体验与考核细节可能略有出入,见谅。
课程信息
开课单位
计算机科学与技术系
总学时
60
理论/实验/实践学时
60/0/0
开课学期
秋、春
评分制
百分制
考核方式
笔试(闭卷)
授课语言
中文
课程简介
算法与理论是计算机科学的核心领域之一。随着大数据时代的来临,传统的算法理论已经不能很好地解决人工智能、物联网、工业制造等领域所遇到的实际问题。本门课程主要介绍基于大数据的新型算法技术,如随机采样、数据降维、数据压缩、分布式计算、流数据计算、聚类、分类、随机优化等,以及相关的理论和数学技巧,如概率计算方法、VC 维、通信复杂度、机器学习理论等。作为一门理论方向课程,本课程旨在帮助学生掌握解决大数据问题所需的理论和算法工具,为相关领域的工程实践打好基础。
English Description
Algorithms and Theory is the core field of computer science. However, as the rapid development of big data era, the traditional algorithmic theory and techniques cannot well handle the practical problems in the areas of Artificial Intelligence, Internet of Things, and Industry. In this course, we aim to introduce the new algorithmic techniques in big data, such as random sampling, dimension reduction, data compression, distributed computing, streaming algorithms, clustering, classification, and stochastic optimization, and the related theory and mathematics, such as probability foundations, VC-dimension, communication complexity, and machine learning theory. As a theory course, it will help students to learn the basic theory and algorithms for handling big data problems, and lay down the foundation for related studies in engineering afterwards.
前置知识涉及的课程
线性代数(B1)、数据结构
课堂概况
课堂授课内容基本被讲义内容包含,不考勤。
到堂听课有三点好处:(1)本门课程讲述的内容基本上还是不易理解的,老师的讲解有助于深化对于知识的把握;(2)老师可能会在课程上给出一些期中期末考点的预告,作为给到堂同学的福利;(3)本课程考试的内容都是上课讲过的内容,讲义中至少1/3的内容是不会讲不会考的,故若旷课可能需要后期与助教更进进度。
老师平易近人,上课时也热心于和学生交流。
教材
Foundations of Data Science, 1st Edition, Avrim Blum, John Hopcroft and Ravindran Kannan, Cambridge University Press 一般来说用不到,上课与考试的主要参考资料是老师的自编讲义
作业
作业量不多,笔者参与的课堂一共有4次作业,建议自己完成,收获很大。
考试情况
课程采用闭卷笔试考核,包含期中与期末。
本课程的考试内容基本上是讲义证明内容的子集,同时难度远低于作业题,对于本门课程的修读,深刻地去理解大数据算法(尤其是其中区别于其他算法的随机性与优化思想)会比硬磕考试更有收获,老师也不会在给分上卡人。
评分细则
课程采用百分制评分。
阿蒙,2026/8/15写于GT-A
最后更新于
