8. GENERACIÓN DE LA PROPUESTA DE MEJORA Y ESTIMACIÓN DE INDICADORES
8.2. PROCESO DE EXTRUSIÓN E INYECCIÓN
8.2.4. Nuevo Rol para el Operario de Extrusión
As we have seen in the Figures 2, 3 and 4 the current visualization of SOM output could be improved for more analytical ability. We introduce a new method to plot SOM output especially designed for large datasets.
Algorithm
1. Filter the results of SOM.
2. Make a polygon with as many sides as the variables in input.
3. Make the radius of the polygon to be the maximum of the value in the dataset.
4. Draw the grid for the polygon.
5. Make segments inside the polygon if the strength of the two variables inside the segment is greater than the specified threshold.
6. Loop Step v for every variable against every other variable
7. Color the segments based on the frequency of variable.
8. Color the line segments based on the threshold of each variable pair plotted.
Figure 3: SOM Visualization in R using the Package
‘Kohonen’
Figure 4: SOM visualization in R using the package
‘SOM’
Source: Infosys Research Source: Infosys Research
Plots
As we can see, this plot is more meaningful than the SOM visualization plots obtained before.
From the figure we can easily deduce that the words ‘free’ and ‘order’ do not have similar relation as ‘credit’ and ‘money’. Understandably so, because if a Spam email is selling something, it will probably have the words ‘order’ and conversely if it is advertising any product or software for ‘free’ download then it wouldn’t have the words ‘order’ in it. High relationship between ‘credit’ and ‘money’ signifies Spam emails advertising for better ‘Credit Score’
programs and other marketing traps.
Figure 6 shows the relationship of each variable-- in this case four popular recurring
words in the Spam database. The number of threads between one variable to another shows the probability of second variable given the first variable. Several threads between ‘free’
and ‘credit’ suggests that Spam emails offering
‘free credit’ (disguised in other forms by fees or deferred interests) are among the most popular.
Using these Spider plots we can analyze several variables at once. This may cause the graph to be messy but sometimes we need to see the complete picture in order to make canonical decisions about the dataset.
From Figure 7 we can see that even though the figure shows 25 variables it is not as cluttered as a Scatter Plot or Bar chart would be if plotted with 25 variables.
Figure 8: Uncolored Representation of Threads in Six variables
Figure 6: SOM visualization in R using Above Algorithm:
Showing Threads, i.e., inter-variable strength) Source: Infosys Research
Figure 5: SOM Visualization in R Using the Above
Algorithm: Showing Segment, i.e., inter-variable dependency Figure 7: Spider Plot showing 25 Sampled Words from the Spam Database
Source: Infosys Research Source: Infosys Research
Source: Infosys Research
Figure 8 shows the different levels of strength between different variables. While
‘contact’ variable is strong with ‘need’ but not enough with ‘help’ it is no surprise that ‘you’
and ‘need’ are strong. Here the idea was only to present the visualization technique and not the analysis of Spam dataset. For more analysis on Spam filtering and Spam analysis one may refer to several independent works on the same [13, 14].
ADVANTAGES
There are several visual and non-visual advantages of using this new plot against the existing plot obtained. This plot has been designed to handle Big data. Most of the existing plots mentioned above are limited in their capacity to scale. Principally if the range of data is large then most of the existing plots tend to get skewed and important information is lost.
By normalizing the data this new plot prevents this issue. By allowing multiple dimensions to be incorporated allows for recognition of indirect relationships.
CONCLUSION
While unstructured data is abundant, free and hidden with information the tools of analyzing the same are still nascent and cost of converting them to structured form is very high. Machine learning is used to classify unstructured data but comes with issues of speed and space constraints. SOM are the fastest machine learning algorithms but their visualization powers are limited. We have presented a naturally intuitive method to visualize SOM outputs which facilitates multi-variable analysis and is also highly scalable.
REFERENCE
1. Grimes, S., Unstructured data and the 80 percent rule. Retrieved from
http://clarabridge.com/default.
aspx?tabid=137.
2. Doan, A., Naughton, J. F., Ramakrishnan, R., Baid, A., Chai, X., Chen, F. and Vuong, B. Q. (2009), Information extraction challenges in managing unstructured data, ACM SIGMOD Record, vol. 37, no.
4, pp. 14-20.
3. Diesner, J., Frantz, T. L. and Carley, K.
M. (2005). Communication networks from the Enron email corpus “It’s always about the people. Enron is no different”.
In Computational & Mathematical Organization Theory, vol. 11, no. 3, pp.
201-228.
4. Chapanond, A., Krishnamoorthy, M.
S., & Yener, B. (2005), Graph theoretic and spectral analysis of Enron email data. In Computational & Mathematical Organization Theory, vol. 11, no.3, pp.
265-281.
5. Peterson, K., Hohensee, M., and Xia, F.
(2011), Email formality in the workplace:
A case study on the enron corpus.
In Proceedings of the Workshop on Languages in Social Media, pp. 86-95.
Association for Computational Linguistics.
6. Buneman, P., Davidson, S., Fernandez, M., and Suciu, D. (1997), Adding structure to unstructured data. Database Theory—ICDT’97, pp. 336-350.
7. Kohonen, T. (1990),The self-organizing map. Proceedings of the IEEE, vol. 78, no. 9, pp. 1464-1480.
8. Waller, N. G., Kaiser, H. A., Illian, J. B., and Manry, M. (1998), A comparison of the classification capabilities of the 1-dimensional kohonen neural network with two pratitioning and three hierarchical cluster analysis algorithms.
Psychometrika, vol. 63, no.1, pp. 5-22.
9. Carpenter, G. A., and Grossberg, S.
(1987), A massively parallel architecture for a self-organizing neural pattern recognition machine. Computer vision, graphics, and image processing, vol. 37, no. 1, pp. 54-115.
10. Kohonen, T., and Somervuo, P. (2002), How to make large self-organizing maps for non-vectorial data. Neural Networks, vol.15, no. 8, pp. 945-952.
11. Wehrens, R & Buydens, L.M.C (2007), Self- and Super-organizing Maps in R: The Kohonen Package. Journal of Statistical Software, vol. 21, no. 5, pp. 1-19.
12. Yan, J. (2012), Self-Organizing Map (with application in gene clustering) in R.
Available at http://cran.r-project.org/
web/packages/som/som.pdf.
13. Dasgupta, A., Gurevich, M., & Punera, K. (2011), Enhanced email spam filtering through combining similarity graphs.
In Proceedings of the fourth ACM international conference on Web search and data mining, pp. 785-794.
14. Cormack, G. V. (2007), Email spam f i l t e r i n g : A s y s t e m a t i c r e v i e w . Foundations and Trends in Information Retrieval, vol. 1, no. 4, pp. 335-455.
NOTES
Index
Automated Content Discovery 48, 49, Big Data
Analytics 4-8, 19, 24, 40-43, 45, 67, Lifecycle 21,
Medical Engine 42- 44 Value, also BDV 27, 29, Campaign Management 31, 32,
Common Warehouse Meta-Model, also CWM 7 Communication Service Providers, also CSPS 27,
Complex Event Processing, also CEP 53-63 Content
Processing Workflows 50
Publishing Lifecycle Management, also CPLM 48,
Management System, also CMS 30, 48, 51 Contingency Funding Planning, also CFP 36, Customer
Dynamics 19-21, 25 Relationship 28, 30
Data Warehouse 4- 5, 30, 38-39, 66, 68 Enterprise Service Bus, also ESB 30 Event Driven
Process Automation
Architecture, also EDA 30-31 Experience Personalization 31
Extreme Content Hub, also ECH 47-51 Global Positioning Service, also GPS 10, 13, 17, 54, 56
Management
Business Process, also BPM 30, Custom Relationship, also CRM 28-30 Information 3, 56-57
Liquidity Risk, also LRM 35-40 Master Data 5-6
Offer 32 Order 30 Retention 31, 32 Metadata
Discovery 6-7 Extractor 50, Governance 6-7 Management 3-8
Net Interest Income Analysis, also NIIA 37 Predictive
Intelligence 19 Modeling 32 Analytics 54
Service Management 31, 33 Supply Chain Planning 9-12, 53 Un-Structured Content Extractor 50 Web Analytics 21
BUSINESS INNOVATION through TECHNOLOGY
Editorial Office: Infosys Labs Briefings, B-19, Infosys Ltd.
Electronics City, Hosur Road, Bangalore 560100, India
Email: [email protected] http://www.infosys.com/infosyslabsbriefings
© Infosys Limited, 2013
Infosys acknowledges the proprietary rights of the trademarks and product names of the other companies mentioned in this issue. The information provided in this document is intended for the sole use of the recipient and for educational purposes only. Infosys makes no express or implied warranties relating to the information contained herein or to any derived results obtained by the recipient from the use of the information in this document. Infosys further does not guarantee the sequence, timeliness, accuracy or completeness of the information and will not be liable in any way to the recipient for any delays, inaccuracies, errors in, or omissions of, any of the information or in the transmission thereof, or for any damages arising therefrom. Opinions and forecasts constitute our judgment at the time of release and are subject to change without notice. This document does not contain information provided to us in confidence by our clients.
Editor Praveen B Malla PhD
Deputy Editor Yogesh Dandawate
Graphics & Web Editor Rakesh Subramanian
Chethana M G Vivek Karkera
IP Manager K V R S Sarma
Marketing Manager Gayatri Hazarika
Online Marketing Sanjay Sahay
Production Manager Sudarshan Kumar V S
Database Manager Ramesh Ramachandran
Distribution Managers Santhosh Shenoy Suresh Kumar V H
How to Reach Us:
Email:
[email protected] Phone: +91 40 44290563
Post:
Infosys Labs Briefings, B-19, Infosys Ltd.
Electronics City, Hosur Road, Bangalore 560100, India
Subscription:
Rights, Permission, Licensing and Reprints:
Infosys Labs Briefings is a journal published by Infosys Labs with the objective of offering fresh perspectives on boardroom business technology.
The publication aims at becoming the most sought after source for thought leading, strategic and experiential insights on business technology management.
Infosys Labs is an important part of Infosys’ commitment to leadership in innovation using technology. Infosys Labs anticipates and assesses the evolution of technology and its impact on businesses and enables Infosys to constantly synthesize what it learns and catalyze technology enabled business transformation and thus assume leadership in providing best of breed solutions to clients across the globe. This is achieved through research supported by state-of-the-art labs and collaboration with industry leaders.
About Infosys
Many of the world’s most successful organizations rely on Infosys to deliver measurable business value. Infosys provides business consulting technology, engineering and outsourcing services to help clients in over 32 countries build tomorrow’s enterprise.
For more information about Infosys (NASDAQ:INFY), visit www.infosys.com
Infosys Labs Briefings
31 %
OF COMPANIES REPORT THEY ARE JUST STARTING TO DEVELOP A MOBILE STRATEGY OR HAVE NO MOBILE STRATEGY AT ALL.