A Text Analytic Approach to Classifying Document Types

Steven Walczak University of South Florida

Abstract

Background: While it is commonly recognized that almost every work and research discipline utilize their own taxonomy, the language used within a specific discipline may also vary depending on numerous factors, including the desired effect of the information being communicated and the intended audience. Different audiences are reached through publication of information, including research results, in different types of publication outlets such as newspapers, newsletters, magazines, websites, and journals. Prior research has shown that students, both undergraduate and graduate, as well as faculty may have a difficult time locating information in different publication outlet types (e.g., magazines, newspapers, journals). The type of publication may affect the ease of understanding and also the confidence placed in the acquired information. A text analytics tool for classifying the source of research as a newsletter (used as a substitute for newspaper articles), a magazine, or an academic journal article has been developed to assist students, faculty, and researchers in identifying the likely source type of information and classifying their own writings with respect to these possible publication outlet types.  Literature Review: Literature on information literacy is discussed as this forms the motivation for the reported research. Additionally, prior research on using text mining and text analytics is examined to better understand the methodology employed, including a review of the original Scale of Theoretical and Applied Research system, adapted for the current research. Research Questions: The primary research question is: Can a text mining and text analytics approach accurately determine the most probable publication source type with respect to being from a newsletter, magazine, or journal? Methodology: A text mining and text analytics algorithm, STAR’ (System for Text Analytics-based Ranking), was developed from a previously researched text mining tool, STAR (Scale of Theoretical and Applied Research), that was used to classify the research type of articles between theoretical and applied research. The new text mining method, STAR’, analyzes the language used in manuscripts to determine the type of publication. This method first mines all words from corresponding publication source types to determine a keyword corpus. The corpus is then used in a text analytics process to classify full newsletters, magazine articles, and journal articles with respect to their publication source. All newsletters, magazine articles, and journal articles are from the library and information sciences (LIS) domain. Results: The STAR’ text analytics method was evaluated as a proof of concept on a specific LIS organizational newsletter, as well as articles from a single LIS magazine and a single LIS journal. STAR’ was able to classify the newsletters, magazine articles, and journal articles with 100% accuracy. Random samples from another similar LIS newsletter and a different LIS journal were also evaluated to examine the robustness of the STAR’ method in the initial proof of concept. Following the positive results of the proof of concept, additional journal, magazine, and newsletter articles were used to evaluate the generalizability of STAR’. The second-round results were very positive for differentiating journals and newsletters from other publication types, but revealed potential issues for distinguishing magazine articles from other types of publications. Discussion: STAR’ demonstrates that the language used for transferring information within a specific discipline does differ significantly depending on the intended recipients of the research knowledge. Further work is needed to examine language usage specific to magazine articles. Conclusions: The STAR’ method may be used by students and faculty to identify the likely source of research or discipline-specific information. This may improve trust in the reliability of information due to different levels of rigor applied to different types of publications. Additionally, the STAR’ classifications may be used by students, faculty, or researchers to determine the most appropriate type of outlet and correspondingly the most appropriate type of audience for the reported information in their own manuscripts, thereby improving the chance for successful sharing of information to appropriate audiences who will deem the information to be reliable, through publication in the most relevant outlet type.

Journal
Journal of Writing Analytics
Published
2017-01-01
DOI
10.37514/jwa-j.2017.1.1.06
CompPile
Open Access
OA PDF Gold
Topics
Export

Citation context

Cited by in this index (1)

  1. Journal of Writing Analytics

References (68)

  1. Aggarwal, C. C., & Zhai, C. X. (2012). An introduction to text mining. In C. C. Aggarwal & C. X. Zhai (Eds.),…
     ↗
  2. Balmas, M. (2014). When fake news becomes real: Combined exposure to multiple news sources and political atti…
     ↗
  3. Baoli, L., Qin, L., & Shiwen, Y. (2004). An adaptive k-nearest neighbor text categorization strategy. ACM Tra…
     ↗
  4. Barclay, D. A. (2017, January 4). The challenge facing libraries in an era of fake news. The Conversation. Re…
     ↗
  5. Bichindaritz, I., & Akkineni, S. (2006). Concept mining for indexing medical literature. Engineering Applicat…
     ↗
Show all 68 →
  1. Bonthron, K., Urquhart, C., Thomas, R., Armstrong, C., Ellis, D., Everitt, J., Fenton, R., Lonsdale, R., McDe…
     ↗
  2. Bragge, J., Thavikulwat, P., & Töyli, J. (2010). Profiling 40 years of research in Simulation & Gaming. Simul…
     ↗
  3. Bui, D. D. A., Del Fiol, G., & Jonnalagadda, S. (2016). PDF text classification to leverage information extra…
     ↗
  4. Calisir, F., & Calisir, F. (2004). The relation of interface usability characteristics, perceived usefulness,…
     ↗
  5. Chiclana, F., Herrera, F., & Herrera-Viedma, E. (2001). Integrating multiplicative preference relations in a …
     ↗
  6. Chu, S. K., Lau, W. W., Chu, D. S., Lee, C. W., & Chan, L. L. (2016). Media awareness among Hong Kong primary…
     ↗
  7. Cockrell, B. J., & Jayne, E. A. (2002). How do I find an article? Insights from a web usability study. The Jo…
     ↗
  8. Cohen, A. M., & Hersh, W. R. (2005). A survey of current work in biomedical text mining. Briefings in Bioinfo…
     ↗
  9. Conway, M. (2010). Mining a corpus of biographical texts using keywords. Literary and Linguistic Computing, 2…
     ↗
  10. Crain, S. P., Zhou, K., Yang, S. H., & Zha, H. (2012). Dimensionality reduction and topic modeling: From late…
     ↗
  11. Davis, P. M., & Cohen, S. A. (2001). The effect of the Web on undergraduate citation behavior 1996-1999. Jour…
     ↗
  12. Dietrich, D., Heller, B., & Yang, B. (2015). Data science and big data analytics. Indianapolis: Wiley.
  13. Fan, W., Wallace, L., Rich, S., & Zhang, Z. (2006). Tapping the power of text mining. Communications of the A…
     ↗
  14. Fitzgerald, M. (2012). Introducing regular expressions. Sebastopol: O'Reilly Media.
  15. Franco, A., Malhotra, N., & Simonovits, G. (2014). Publication bias in the social sciences: Unlocking the fil…
     ↗
  16. Frigui, H., & Nasraoui, O. (2004). Simultaneous clustering and dynamic keyword weighting for text documents. …
     ↗
  17. Gandomi, A., & Haider, M. (2015). Beyond the hype: Big data concepts, methods, and analytics. International J…
     ↗
  18. Goutte, C. (2008). A probabilistic model for fast and confident categorisation of textual documents. In M. W.…
     ↗
  19. Griffiths, J. R., & Brophy, P. (2005). Student searching behavior and the Web: Use of academic resources and …
  20. Head, A. J., & Wihbey, J. (2014, July 17). At sea in a deluge of data. The Chronicle of Higher Education, 3. …
  21. Herrera, F., & Herrera-Viedma, E. (2000). Linguistic decision analysis: Steps for solving decision problems u…
     ↗
  22. Hotho, A., Nürnberger, A., & Paaß, G. (2005). A brief survey of text mining. LDV Forum, 20(1), 19-62. Retriev…
     ↗
  23. Hovland, C. I., & Weiss, W. (1951). The influence of source credibility on communication effectiveness. Publi…
     ↗
  24. Iriondo, I., Planet, S., Socoró, J. C., Martínez, E., Alías, F., & Monzo, C. (2009). Automatic refinement of …
     ↗
  25. Isa, D., Kallimani, V. P., & Lee, L. H. (2009). Using the self organizing map for clustering of text document…
     ↗
  26. Jagadish, H. V. (2008). The conference reviewing crisis and a proposed solution. ACM SIGMOD Record, 37(3), 40…
     ↗
  27. Jamieson, S. (2016). What the citation project tells us about information literacy in college composition. In…
     ↗
  28. Kellogg, D. L., & Walczak, S. (2007). Nurse scheduling: From academia to implementation or not? Interfaces, 3…
     ↗
  29. Kim, S. B., Han, K. S., Rim, H. C., & Myaeng, S. H. (2006). Some effective techniques for Naive Bayes text cl…
     ↗
  30. Kimble, J. (2013). You think the law requires legalese. Michigan Bar Journal, 92(11), 48-50. Retrieved from h…
  31. Kissel, F., Wininger, M. R., Weeden, S. R., Wittberg, P. A., Halverson, R. S., Lacy, M., & Huisman, R. K. (20…
     ↗
  32. Klosterman, M. L., Sadler, T. D., & Brown, J. (2012). Science teachers' use of mass media to address socio-sc…
     ↗
  33. Korhonen, A., Séaghdha, D. O., Silins, I., Sun, L., Högberg, J., & Stenius, U. (2012). Text mining for litera…
     ↗
  34. Kumar, M. N. (2014). Review of the ethics and etiquettes of time management of manuscript peer review. Journa…
     ↗
  35. Landers, R. N., & Callan, R. C. (2011). Casual social games as serious games: The psychology of gamification …
     ↗
  36. Laskin, M., & Haller, C. R. (2016). Up the mountain without a trail: Helping students use source networks to …
     ↗
  37. Laubersheimer, J., Ryan, D., & Champaign, J. (2016). InfoSkills2Go: Using badges and gamification to teach in…
     ↗
  38. Lewis, S. C. (2008). Where young adults intend to get news in five years. Newspaper Research Journal, 29(4), 36-52.
     ↗
  39. Lloyd, A. (2005). Information literacy: different contexts, different concepts, different truths? Journal of …
     ↗
  40. Luchins, D. J. (2007). Corporate speak and the psychiatric profession. Administration and Policy in Mental He…
     ↗
  41. Ma, J., Xu, W., Sun, Y. H., Turban, E., Wang, S., & Liu, O. (2012). An ontology- based text-mining method to …
     ↗
  42. McCune, J. C. (1999). Do you speak computerese? Management Review, 88(2), 10-12. Retrieved from http://search…
  43. Metzger, M. J., Flanagin, A. J., & Zwarun, L. (2003). College student Web use, perceptions of information cre…
     ↗
  44. Mogge, D. (1999). Seven years of tracking electronic publishing: The ARL Directory of Electronic Journals, Ne…
     ↗
  45. Polites, G. L., & Watson, R. T. (2008). The centrality and prestige of CACM. Communications of the ACM, 51(1)…
     ↗
  46. Reed, K. L. (1999). Mapping the literature of occupational therapy. Bulletin of the Medical Library Associati…
  47. Robertson, S. (2004). Understanding inverse document frequency: On theoretical arguments for IDF. Journal of …
     ↗
  48. Senellart, P. P., & Blondel, V. D. (2004). Automatic discovery of similar words. In M. W. Berry (Ed.), Survey…
     ↗
  49. Stiller, J., Hartmann, S., Mathesius, S., Straube, P., Tiemann, R., Nordmeier, V., Krüger, D., & Upmeier zu B…
     ↗
  50. Sun, Y., Deng, H., & Han, J. (2012). Probabilistic models for text mining. In C. C. Aggarwal & C. X. Zhai (Ed…
     ↗
  51. Susman, G. I, & Evered, R. D. (1978). An assessment of the scientific merits of action research. Administrati…
     ↗
  52. Teddlie, C., & Yu, F. (2007). Mixed methods sampling: A typology with examples. Journal of Mixed Methods Rese…
     ↗
  53. Teranes, P. S. (2013). Make it as simple as you can. Michigan Bar Journal, 92(11), 52-53. Retrieved from http…
  54. Ting, S. L., Ip, W. H., & Tsang, A. H. (2011). Is Naive Bayes a good classifier for document classification. …
  55. Tseng, Y. H., Chang, C. Y., Rundgren, S. N. C., & Rundgren, C. J. (2010). Mining concept maps from news stori…
     ↗
  56. Walczak, S., & Kellogg, D.L. (2015). A heuristic text analytic approach for classifying research articles. In…
     ↗
  57. Weiss, S. M., Apte, C., Damerau, F. J., Johnson, D. E., Oles, F. J., Goetz, T., & Hampp, T. (1999). Maximizin…
     ↗
  58. Wiebe, T. (2016). The information literacy imperative in higher education. Liberal Education, 102(1), 52-55. …
  59. Worsley, A. (1989). Perceived reliability of sources of health information. Health Education Research, 4(3), …
     ↗
  60. Yamamoto, M., & Church, K. W. (2001). Using suffix arrays to compute term frequency and document frequency fo…
     ↗
  61. Yancey, K. B. (2016). Creating and exploring new worlds: Web 2.0 information literacy and the ways we know. I…
     ↗
  62. Yao, Q. (2009). An evidence of frame building: Analyzing the correlations among the frames in Sierra Club new…
     ↗
  63. Young, M. E., Norman, G. R., and Humphreys, K. R. (2008). The role of medical language in changing public per…
     ↗