Select your font size 
 
about us products & services consulting & support news & events contact us
Centuries-old techniques developed by Thomas Bayes find modern applications because they are simple and effective.

Bayesian Techniques Assist Automated Decision Tools - Louisiana

print this article 
 

Centuries-old techniques developed by Thomas Bayes find modern applications because they are simple and effective.


(1) Bayesian techniques are employed in the automated filtering of unwanted spam, the formation of medical diagnoses, the detection of viruses, and in several other ways that advance compatible business objectives.

Bayes Theorem uses Conditional Probability to calculate the probability of A given B, provided that the probability of B given A, the probability of A and the probability of B are all known.

For example, suppose an automated program could determine that a particular phrase is present in 70% of spam and 50% of non-spam emails, and that an email is 90% likely to be spam, and suppose that A means "the email is spam", B means "the email is not spam", and C means "the phrase is present". Then P(A), the probability of A, is 90%. P(B) is 10%. P(C|A), the probability of C given that A is true, is 70%. P(C|B) is 50%. P(C) = (0.70 * 0.90) + (0.50 * 0.10) = 0.68 = 68%. Some useful probabilities for classifying the email would then be P(A|C), the probability that the email is spam given that the phrase is present, and P(B|C), the probability that the email is not spam given that the phrase is present. Using conditional probability, P(A|C) = P(A) * P(C|A) / P(C). P(A|C) = 0.90 * 0.70 / 0.68 ~= 0.92647 ~= 93%. P(B|C) = P(B) * P(C|B) / P(C) = 0.10 * 0.50 / 0.68 ~= 0.07353 ~= 7%. Therefore, the probability that the email is spam, based on that one data point, is 93%, and the probability that it is not spam is 7%.

To improve the accuracy of this technique, a computer program could analyze thousands of data points in less than a tenth of a second, which is approximately how long it takes to download an email message. The program could test for phrase D, phrase E, phrase F, etc., and use data about each one to modify the overall confidence that the email is spam or not spam. In hand-waving mathematical terms, if C, D and E are found to be true, the program can automatically determine P(A|C,D,E) (the probability that the email is spam given that C, D, and E are true) using an extended version of Bayes Theorem, applied to a Bayesian network. By using good tokens (i.e. by asking the right questions, using all available information), this technique can be up to 99.5% accurate with 0.3% false positives.

A related algorithm is state-based, so that rather than using a directed acyclic graph of conditional probabilities (a Bayesian network), one could automatically produce a set of known states, along with transition matrices (stochastic matrices) showing the probability of moving from a given state to another (transition probabilities). These stochastic matrices fall naturally from statistics coming from a large enough set. CRM114 uses this technique to improve spam detection accuracy beyond naive Bayesian techniques, and suggests innovative approaches to potentially bring 99.999% reliability.

The applications of statistical techniques such as described above go beyond spam detection. For instance, one could imagine the same or similar techniques being used in the medical field to classify DNA, in industry to properly direct calls that come through automated call systems, in elections to predict how a particular message may affect polls, etc.

To learn more about how statistical methods can be used to improve your business, please contact us.

Works Cited

  1. J J O'Connor and E F Robertson. "Thomas Bayes." From The MacTutor History of Mathematics archive.
  2. William S. Yerazunis. "The Spam Filtering Plateau at 99.9% Accuracy and How to Get Past It." From
  3. CRM114 - the Controllable Regex Mutilator.
  4. Eric W. Weisstein. "Bayes' Theorem." From MathWorld--A Wolfram Web Resource.
  5. Eric W. Weisstein. "Conditional Probability." From MathWorld--A Wolfram Web Resource.
  6. Paul Graham. "Better Bayesian Filtering." From
  7. Paul Graham.

Most Recent Website and Regional Updates

 Transparen Toronto Office Locations
Addresses of Transparen Corporation offices in Toronto.

 
 Emergency Management Services
The prototypical emergency involves a shutdown of essential services for a finite period of time. What will your organization do when a world-wide financial crisis strikes?

 
 High Scalability - Large Systems Optimization
Transparen Corporation lends its expertise to clients experiencing rapid and sudden growth in traffic or server utilization, bottlenecks, systems instability, downtime during peak traffic, or which would like to plan to avoid such issues.

 
 Fast RAID Server Data Recovery Service
Transparen's Vancouver International Response Team provides the option in Canada and USA to get a raid server back running in hours - eliminating costly waiting associated with typical RAID recoveries.

 
 Data Recovery Service
Have you deleted a mission critical file? Accidentally dropped a computer, or formatted a hard drive? No recent backup? Mistakes can happen, but the data might still be there.

 
 About Transparen
Transparen is committed to serving its clients.

 
 Research Tools
Measure human resource allocation and collect data with the goal of determining patterns that will bring forward actionable insights which may lead to policy changes, saving money and improving quality of service.

 

Google
 
Web transparen.com

Contact Information

Related Information

   
 
E C M | © 2003-2007 Transparen Corp.      

Standardized Services: Data Recovery Service / Creative Services / Premium Web Hosting Services / System Administration Tech Support Services
Recent Projects: Full-Service Mortgage and Financing Company / System to manage flights from Vancouver to Tofino / Photo exchange verification service
Our Vancouver BC Server Proudly Hosts: automated parking and revenue control systems, leafside lane at southlands, cost effective alternative power sources, Higher Grade Learning Centres, pacific forage bag supply, sunburst medical, neosonic design, roger mahler photography - passionate, intriguing, desirable, the connection between east and west, affordable flights to victoria and tofino, low interest mortgage brokers in vancouver, richmond, surrey, toronto, Toronto Calgary and Vancouver IT staffing and talent search
* Abbeville * Alexandria * Baker * Bastrop * Baton Rouge * Bogalusa * Bossier City * Breaux Bridge * Bunkie * Carencro * Covington * Crowley * Denham Springs * DeQuincy * De Ridder * Donaldsonville * Eunice * Franklin * Gonzales * Grambling * Gretna * Hammond * Harahan * Houma * Jeanerette * Jennings * Kaplan * Kenner * Lacombe * Lafayette * Lake Charles * Leesville * Mandeville * Mansfield * Marksville * Minden * Monroe * Morgan City * Natchitoches * New Iberia * New Orleans * New Roads * Oakdale * Opelousas * Patterson * Pineville * Plaquemine * Ponchatoula * Port Allen * Rayne * Ruston * St. Gabriel * St. Martinville * Scott * Shreveport * Slidell * Springhill * Sulphur * Tallulah * Thibodaux * Ville Platte * Westlake * West Monroe * Westwego * Winnfield * Winnsboro * Zachary * Abita Springs * Addis * Amite City * Arcadia * Arnaudville * Baldwin * Ball * Basile * Benton * Bernice * Berwick * Blanchard * Boyce * Broussard * Brusly * Campti * Chatham * Cheneyville * Church Point * Clayton * Clinton * Colfax * Columbia * Cottonport * Cotton Valley * Coushatta * Cullen * Delcambre * Delhi * Dubach * Duson * Elizabeth * Elton * Erath * Eros * Evergreen * Farmerville * Ferriday * Fordoche * Franklinton * Gibsland * Glenmora * Golden Meadow * Gramercy * Grand Coteau * Grand Isle * Greensburg * Greenwood * Gueydan * Haughton * Haynesville * Henderson * Homer * Hornbeck * Independence * Iota * Iowa * Jackson * Jean Lafitte * Jena * Jonesboro * Jonesville * Keachi * Kentwood * Kinder * Krotz Springs * Lake Arthur * Lake Providence * Lecompte * Leonville * Livingston * Livonia * Lockport * Logansport * Lutcher * Madisonville * Mamou * Mangham * Mansura * Many * Maringouin * Marion * Melville * Montgomery * Mooringsport * Mount Lebanon * Newellton * New Llano * Oak Grove * Oberlin * Oil City * Olla * Pearl River * Plain Dealing * Pollock * Port Barre * Rayville * Richwood * Ridgecrest * Ringgold * Roseland * Rosepine * St. Francisville * St. Joseph * Sarepta * Sibley * Simmesport * Slaughter * Sorrento * Springfield * Sterlington * Stonewall * Sunset * Tullos * Urania * Vidalia * Vienna * Vinton * Vivian * Walker * Washington * Waterproof * Welsh * White Castle * Wisner * Woodworth * Youngsville * Zwolle * Albany * Anacoco * Angie * Ashland * Athens * Atlanta * Baskin * Belcher * Bienville * Bonita * Bryceland * Calvin * Cankton * Castor * Chataignier * Choudrant * Clarence * Clarks * Collinston * Converse * Delta * Dixie Inn * Dodson * Downsville * Doyline * Dry Prong * Dubberly * East Hodge * Edgefield * Epps * Estherwood * Fenton * Fisher * Florien * Folsom * Forest * Forest Hill * French Settlement * Georgetown * Gilbert * Gilliam * Goldonna * Grand Cane * Grayson * Grosse Tete * Hall Summit * Harrisonburg * Heflin * Hessmer * Hodge * Hosston * Ida * Jamestown * Johnson's Bayou * Junction City * Kilbourne * Killian * Lillie * Lisbon * Longstreet * Loreauville * Lucky * McNary * Martin * Maurice * Mermentau * Mer Rouge * Montpelier * Moreauville * Morganza * Morse * Mound * Napoleonville * Natchez * Noble * North Hodge * Norwood * Oak Ridge * Palmetto * Parks * Pilot Town * Pine Prairie * Pioneer * Plaucheville * Pleasant Hill * Port Fourchon * Port Vincent * Powhatan * Provencal * Quitman * Reeves * Richmond * Robeline * Rodessa * Rosedale * Saline * Shongaloo * Sicily Island * Sikes * Simpson * Simsboro * South Mansfield * Spearsville * Stanley * Sun * Tangipahoa * Tickfaw * Turkey Creek * Varnado * Wilson