Top > Info > Data Mining > 1-1-1. What is data mining?
▷▶ What is data mining?
[DM의 정의] : DM의 정의는 정의자의 관점과 background에 의존한다. DM의 논문으로부터 출처한 몇 가지 정의가 있다.
- Fayyad : DM은 근거가 확실하고, 기발하고. 잠재적으로 유용하고,
최종적으로는 데이터의 패턴을 이해하는 것을 식별하는 특별한 과정이다.
- Zekulin : DM은 큰 데이터베이스로부터 알 수 없고 이해할 수 있고, 활용 가능한
정보를 뽑아내는 과정이고, 그것을 중대한 사업결정을 내리는데 이용하는 과정이다.
- Ferruzza : DM은 데이터 안에서 알 수 없는 관계나 패턴을 구별짓기 위한 지식
발견과정에 사용되어진 방법들의 집합이다.
- John : DM은 데이터에서 유리한 패턴을 발견하는 과정이다.
- Parsaye : DM은 대용량 데이터 베이스에서 보여지는 알 수 없고 기대할 수 없는
정보의 패턴에 대한 것을 알아서 결정을 support하는 과정이다.
[DM이란]
· Decision Trees
· Neural Networks
· Rule Induction
· Nearest Neighbors
· Gentic Algorithms-Mehta
[하드웨어와 소프트웨어]
· 하드웨어 제조자들이 빠르게 DM에 관련된 계산할 수 있는 자격조건을
강조함.
· 소프트웨어 제공자들은 경쟁의 측면을 강조한다.
· Database들은
매우 만드는데도 비싸고 유지하는데도 돈이 들어 감. 상대적으로 적은 투자를 했을
때, DM의 tool들은 데이터 안에 숨겨진 정보의 높은 이윤을 줄 수 있는 neggets을
발견하는 것을 제공.
· 공급업체의 목적 : 하드웨어와 소프트웨어의 공급업체들은
시장이 포화상태가 되기 전에 DM의 상품을 팔아서 자본을 만드는 것
[현재의 DM의 상품]
- IBM : "Intelligent Miner"
- Tandem : "Relational
Data Miner"
- Angoss Software : "KnowledgeSEEKER"
- Thinking
Machines Corporation : "DarwinTM"
- NeoVista Software : "ASIC"
- ILS Decision Systems, Inc. : "Clementine"
- DataMind Corporation
: "DataMind Data Cruncher"
- Silicon Graphics : "MineSet"
- California Scientific Software : "BrainMaker"
- WizSoft Corporation
: "WizWhy"
- Lockheed Corporation : "Recon"
- SAS
Corporation : "SAS Enterprise Miner"
· 포괄적인 패키지 외에도 특정 분야에 대한 목적으로 만들어진 상품들이
다양함.
· 컴퓨터 과학자와 통계학자사이의 다른점 : 통계학자→아이디어→논문,
컴퓨터 과학자→아이디어→회사.
[현재 DM의 상품들의 특징]
○ Attractive GUI to : · Data bases (query language), · Suite
of data analysis procedures.
○ Windows style interface : · Flexible convenient
input
- point and click icons and menus
- input dialog boxes
- diagrams to describe analyses
- sophisticated graphical views of the
output
- a variety of data plots
- slick graphical representations
: trees, networks, flight simulation, etc.
○ Convenient manipulation of the results.
· 일반적으로 패키지들을
DM의 전문가들과 같이 decision maker들이 정한다.
[DM 패키지에 의해서 제공되어지는 통계적 분석 과정들]
· Decision tree induction (C4.5, CART, CHAID)
· Rule induction
(AQ, CN2, Recon, etc.)
· Nearest neighbors ("case based reasoning")
· Clustering methods ("data segmentation")
· Association rules("market
basket analysis")
· Feature extraction
· Visualization ·Neural
networks
· Bayesian belief networks("graphical models")
·
Genetic algorithms
· Self-organizing maps
· Neuro-fuzzy systems
[DM 패키지에서 제공하지 않는 것]
· Hypothesis testing
· Logistic regression
· Experimental design
· GLM
· Response surface modeling
· Canonical correlation
· ANOVA, MANOVA, etc.
· Principal components
· Linear regression
· Factor analysis
· Discriminant analysis
[ Back to the Page ]
Top > Info > Data Mining > 1-1-1. What is data mining?