Showing posts sorted by relevance for query learning. Sort by date Show all posts
Showing posts sorted by relevance for query learning. Sort by date Show all posts

Friday, January 09, 2015

Machine Learning ebooks

Free ML ebooks:

1. The LION Way: Machine Learning plus Intelligent Optimization

Author/s: Roberto Battiti, Mauro Brunato
Publisher: Lionsolver, Inc., 2013

Learning and Intelligent Optimization (LION) is the combination of learning from data and optimization applied to solve complex problems. This book is about increasing the automation level and connecting data directly to decisions and actions.

2. A Course in Machine Learning

Author/s: Hal Daumé III
Publisher: ciml.info, 2012

This is a set of introductory materials that covers most major aspects of modern machine learning (supervised and unsupervised learning, large margin methods, probabilistic modeling, etc.). It's focus is on broad applications with a rigorous backbone.

3. A First Encounter with Machine Learning

Author/s: Max Welling
Publisher: University of California Irvine, 2011

The book you see before you is meant for those starting out in the field of machine learning, who need a simple, intuitive explanation of some of the most useful algorithms that our field has to offer. A prelude to the more advanced text books.

4. Bayesian Reasoning and Machine Learning

Author/s: David Barber
Publisher: Cambridge University Press, 2011

The book is designed for final-year undergraduate students with limited background in linear algebra and calculus. Comprehensive and coherent, it develops everything from basics to advanced techniques within the framework of graphical models.

5. Introduction to Machine Learning

Author/s: Amnon Shashua
Publisher: arXiv, 2009

Introduction to Machine learning covering Statistical Inference (Bayes, EM, ML/MaxEnt duality), algebraic and spectral methods (PCA, LDA, CCA, Clustering), and PAC learning (the Formal model, VC dimension, Double Sampling theorem).

6. The Elements of Statistical Learning: Data Mining, Inference, and Prediction

Author/s: T. Hastie, R. Tibshirani, J. Friedman - Springer, 2009
This book brings together many of the important new ideas in learning, and explains them in a statistical framework. The authors emphasize the methods and their conceptual underpinnings rather than their theoretical properties.

7. Reinforcement Learning

Author/s: C. Weber, M. Elshaw, N. M. Mayer
Publisher: InTech, 2008

This book describes and extends the scope of reinforcement learning. It also shows that there is already wide usage in numerous fields. Reinforcement learning can tackle control tasks that are too complex for traditional controllers.

8. Machine Learning

Author/s: Abdelhamid Mellouk, Abdennacer Chebira
Publisher: InTech, 2009

Neural machine learning approaches, Hamiltonian neural networks, similarity discriminant analysis, machine learning methods for spoken dialogue simulation and optimization, linear subspace learning for facial expression analysis, and more.

9. Reinforcement Learning: An Introduction

Author/s: Richard S. Sutton, Andrew G. Barto
Publisher: The MIT Press, 1998

The book provides a clear and simple account of the key ideas and algorithms of reinforcement learning. It covers the history and the most recent developments and applications. The only necessary mathematical background are concepts of probability.

10. Gaussian Processes for Machine Learning

Author/s: Carl E. Rasmussen, Christopher K. I. Williams
Publisher: The MIT Press, 2005

Gaussian processes provide a principled, practical, probabilistic approach to learning in kernel machines. The treatment is comprehensive and self-contained, targeted at researchers and students in machine learning and applied statistics.

11. Machine Learning, Neural and Statistical Classification

Author/s: D. Michie, D. J. Spiegelhalter
Publisher: Ellis Horwood, 1994

The book provides a review of different approaches to classification, compares their performance on challenging data-sets, and draws conclusions on their applicability to realistic industrial problems. A wide variety of approaches has been taken.

12. Introduction To Machine Learning

Author/s: Nils J Nilsson, 1997
This book concentrates on the important ideas in machine learning, to give the reader sufficient preparation to make the extensive literature on machine learning accessible. The author surveys the important topics in machine learning circa 1996.

13. Inductive Logic Programming: Techniques and Applications

Author/s: Nada Lavrac, Saso Dzeroski
Publisher: Prentice Hall, 1994

This book is an introduction to inductive logic programming. It covers empirical inductive logic programming with applications in knowledge acquisition, inductive program synthesis, inductive data engineering, and knowledge discovery in databases.

14. Practical Artificial Intelligence Programming in Java

Author/s: Mark Watson
Publisher: Lulu.com, 2008

The book uses the author's libraries and the best of open source software to introduce AI (Artificial Intelligence) technologies like neural networks, genetic algorithms, expert systems, machine learning, and NLP (natural language processing).

15. Information Theory, Inference, and Learning Algorithms

Author/s: David J. C. MacKay
Publisher: Cambridge University Press, 2003

A textbook on information theory, Bayesian inference and learning algorithms, useful for undergraduates and postgraduates students, and as a reference for researchers. Essential reading for students of electrical engineering and computer science.


Friday, November 21, 2014

ML books

Thursday, December 21, 2023

On Audit and Certification of Machine Learning Systems

Obviously, machine learning applications are being used more and more in a wide variety of fields. The general rule today is that in the absence of analytical models, one always turns to machine learning. In itself, machine learning has become synonymous with artificial intelligence. The reverse is also true - artificial intelligence today is machine learning. Sometimes this definition is somewhat limited, and they only talk about artificial neural networks and deep learning in the context of artificial intelligence, but this does not change the essence of the matter. At the same time, it is also obvious that the spread of machine learning technologies leads to the need for their application in the so-called critical areas, where there are special requirements for confirming the operability and quality of software. These areas include, for example, avionics, nuclear power, autonomous vehicles, etc. Audit and, of course, certification are the procedures for evaluating machine learning models. - from our new paper

Sunday, May 01, 2022

A Survey of Adversarial Attacks and Defenses for image data on Deep Learning

This article provides a detailed survey of the so-called adversarial attacks and defenses. These are special modifications to the input data of machine learning systems that are designed to cause machine learning systems to work incorrectly. The article discusses traditional approaches when the problem of constructing adversarial examples is considered as an optimization problem - the search for the minimum possible modifications of correlative data that ”deceive” the machine learning system. As tasks (goals) for adversarial attacks, classification systems are almost always considered. This corresponds, in practice, to the so-called critical systems (driverless vehicles, avionics, special applications, etc.). Attacks on such systems are obviously the most dangerous. In general, sensitivity to attacks means the lack of robustness of the machine (deep) learning system. It is robustness problems that are the main obstacle to the introduction of machine learning in the management of critical systems. - from our new paper

Friday, December 28, 2012

Data Science books

Free books on Data mining, statistics and related areas:


Introduction to Information Retrieval, by Manning, Raghavan, and Schütze

A first encounter with machine learning by Welling

Gaussian processes for Machine Learning by C.E. Rasmussen

The Elements of Statistical Learning, by Hastie, Tibshirani, and Friedman

Introduction to Machine Learning by Smola, Vishwanathan

Think Bayes by Downey

Mining of Massive datasets by Rajamaran, Leskovic, and Ullman

Bayesian Reasoning and Machine Learning by D. Barber

Information Theory, Inference, and Learning Algorithms by D.Mackay

Foundations of Statistical Natural Language Processing by Manning and Schütze

Data Jujitsu by D.J. Patil

Building Data Science Teams by D.J. Patil

Network Science by A.-L. Barabasi.

Comments and new titles are more than welcome. I would like to collect this list for my students.

Monday, March 25, 2024

On Real-Time Model Inversion Attacks Detection

The article deals with the issues of detecting adversarial attacks on machine learning models. In the most general case, adversarial attacks are special data changes at one of the stages of the machine learning pipeline, which are designed to either prevent the operation of the machine learning system, or, conversely, achieve the desired result for the attacker. In addition to the well-known poisoning and evasion attacks, there are also forms of attacks aimed at extracting sensitive information from machine learning models. These include model inversion attacks. These types of attacks pose a threat to machine learning as a service (MLaaS). Machine learning models accumulate a lot of redundant information during training, and the possibility of revealing this data when using the model can become a serious problem.

from our new paper

Wednesday, February 29, 2012

Machine Learning

Topics from Stanford's Machine Learing class:

Supervised learnings.

In supervised learning, one has a set of data with features and labels.

Linear Regression – one/multiple variables
Gradient Descent - a general algorithm for minimizing a function
Logistic Regression – This is useful when predicting classification type results. For example, are you looking for a yes or no result. Does the patient have cancer? Will the customer buy my new product? It can also be helpful for more than 2 results. What color will a person choose (red, blue, green, silver)?
Neural Networks – A learning algorithm that is modeled after the brain. Think of neurons.

Unsupervised Learning

In unsupervised learning, one has a set of data with no features and labels. Can some structure be found for the data?

Clustering – The most popular technique is K-means.
PCA (Principal Components Analysis) – speed up a learning algorithm

Anomaly Detection

This section covers methods to determine if data is bad. Bad data is considered an anomaly.

Recommender Systems

Like the name says, recommender systems are used to make recommendations. Companies like Netflix use recommender systems to recommend new movies to customers. LinkedIn also recommends people to connect with. This is a fairly hot topic in the tech world right now.

Content Based(Features)
  Modified Linear Regression
Non-content Based(No Features)
  Collaborative Filtering
  Matrix Factorization

/via Data Science

Thursday, June 06, 2024

On Certification of Artificial Intelligence Systems

Machine learning systems are today the main examples of the use of Artificial Intelligence in a wide variety of areas. From a practical point of view, we can say that machine learning is synonymous with the concept of Artificial Intelligence. The spread of machine learning technologies leads to the need for their application in the so-called critical areas: avionics, nuclear energy, automatic driving, etc. Traditional software, for example, in avionics, undergoes special certification procedures that cannot be directly transferred to machine learning models. The article discusses approaches to the certification of machine learning models. - from our new paper On Certification of Artificial Intelligence Systems

Tuesday, May 28, 2013

Deep Learning

This tutorial will teach you the main ideas of Unsupervised Feature Learning and Deep Learning. By working through it, you will also get to implement several feature learning/deep learning algorithms, get to see them work for yourself, and learn how to apply/adapt these ideas to new problems - from Stanford tutorial

Wednesday, September 25, 2013

How do you explain Machine Learning?

It looks like the marketing won in math too. As per this thread How do you explain Machine Learning and Data Mining to non Computer Science people? any statistics is machine learning. As I can guess, everything unknown is a deep learning ...

Tuesday, May 03, 2022

On a formal verification of machine learning systems

The paper deals with the issues of formal verification of machine learning systems. With the growth of the introduction of systems based on machine learning in the so-called critical systems (systems with a very high cost of erroneous decisions and actions), the demand for confirmation of the stability of such systems is growing. How will the built machine learning system perform on data that is different from the set on which it was trained? Is it possible to somehow verify or even prove that the behavior of the system, which was demonstrated on the initial dataset, will always remain so? There are different ways to try to do this. The article provides an overview of existing approaches to formal verification. All the considered approaches already have practical applications, but the main question that remains open is scaling. How applicable are these approaches to modern networks with millions and even billions of parameters? - from our new paper

Wednesday, June 05, 2013

Deep Learning with SVM

An interesting paper: Deep Learning using Support Vector Machines To date, deep learning for classi cation using fully connected layers and convolutional layers have almost always used softmax layer objective to learn the lower level parameters. There are exceptions, notably the supervised embedding with nonlinear NCA, and semisupervised deep embedding. In this paper, we propose using multiclass SVM's objective to train deep neural nets for classi fication tasks.

Friday, April 11, 2025

Large Language Models in Cyberattacks

The article provides an overview of the practice of using large language models (LLMs) in cyberat-tacks. Artificial intelligence models (machine learning and deep learning) are applied across various fields,with cybersecurity being no exception. One aspect of this usage is offensive artificial intelligence, specificallyin relation to LLMs. Generative models, including LLMs, have been utilized in cybersecurity for some time,primarily for generating adversarial attacks on machine learning models. The analysis focuses on how LLMs,such as ChatGPT, can be exploited by malicious actors to automate the creation of phishing emails and mal-ware, significantly simplifying and accelerating the process of conducting cyberattacks. Key aspects of LLMusage are examined, including text generation for social engineering attacks and the creation of maliciouscode. The article is aimed at cybersecurity professionals, researchers, and LLM developers, providing themwith insights into the risks associated with the malicious use of these technologies and recommendations forpreventing their exploitation as cyber weapons. The research emphasizes the importance of recognizingpotential threats and the need for active countermeasures against automated cyberattacks. - from our new paper

Tuesday, April 05, 2016

On data sharing in education

Yousef Ibrahim Daradkeh, Mujahed ALdhaifallah, Dmitry Namiot "On Data Sharing in Educational Processes", International Journal of Emerging Technologies in Learning, vol. 11, no. 4, 2016, pp. 28-31

In this paper, we present one model for data sharing in educational classes. Typically, Learning Management Systems present data stores for keeping educational materials as well as the conversations between teachers and students. In our model, we propose a peer to peer data exchange via smartphones. With the high penetration of smartphones across students, the ability to support one-to-one communication with teachers could be a good add-on for the traditional learning support systems. This ability could be especially useful for on-demand organized classes, where the standard support is very costly.

Sunday, September 20, 2015

To Deep or not to Deep

An interesting discussion: Will deep learning make other Machine Learning algorithms obsolete?

Our own answer - No. Very often, simpler algorithms like logistic regression will work fine. Deep Learning success depends on the data volume. It should be huge and it is not always true.

Sunday, October 12, 2025

On Image Augmentation

The paper considers methods of natural image augmentation, i.e. those whose application results are close to natural impacts on environmental objects that machine learning models may encounter in industrial applications: the influence of weather conditions; operating features or malfunctions of device cameras, etc. The paper presents a taxonomy of methods of natural image augmentation, which includes weather artifacts, camera artifacts, and background substitution for the main object in the image. Existing software libraries for image augmentation are considered in detail, and their shortcomings and limitations are described. The architecture and implementation of a new open library for image augmentation are presented, and the results of its testing on specialized datasets are given. - On Natural Image Augmentation to Increase Robustness of Machine Learning Models

Thursday, September 09, 2010

Friday, December 01, 2023

Certification & audit for machine learning systems

Presentation on audit of machine learning systems. Auditing should be a mandatory procedure for industrial AI systems.

Tuesday, January 08, 2019

Machine learning in software development

The subject of the article is the “coding style” concept and the main approaches to detecting the individual style of a programmer. The entire process of creating a software product from this point of view and the main features of programming style are analyzed. It emphasizes the relevance and commercial significance of the problem in terms of product support, plagiarism, work of a large developer’s community in a single repository, an evolution of developer skills. Computational stylometry issues, a possibility of using programming paradigms as an additional factor of style identification are considered. It offers the idea of creating a software tool that allows to identify the style of the author who wrote a particular program fragment and allows less experienced developers to follow the rules accepted in the major part of the repository and determined by coding style of "experts", which leads the code to a uniform format that is easier to maintain and make adjustments. Globally, this stage of analyzing the original (and then the modified code) allows improving the existing algorithms for automatic synthesis of programs.

Our new paper: Using Machine Learning Methods to Establish Program Authorship