LATIN 2006: Theoretical Informatics: 7th Latin American Symposium, Valdivia, Chile, March 20-24, 2006, Proceedings (Lecture Notes in Computer Science, 3887) 9783540327554, 354032755X

This book constitutes the refereed proceedings of the 7th International Symposium, Latin American Theoretical Informatic

179 62 8MB

English Pages 830 [828] Year 2006

Report DMCA / Copyright

DOWNLOAD PDF FILE

Table of contents :
Frontmatter
Keynotes
Algorithmic Challenges in Web Search Engines
RNA Molecules: Glimpses Through an Algorithmic Lens
Squares
Matching Based Augmentations for Approximating Connectivity Problems
Modelling Errors and Recovery for Communication
Lossless Data Compression Via Error Correction
The Power and Weakness of Randomness in Computation
Regular Contributions
A New GCD Algorithm for Quadratic Number Rings with Unique Factorization
On Clusters in Markov Chains
An Architecture for Provably Secure Computation
Scoring Matrices That Induce Metrics on Sequences
Data Structures for Halfplane Proximity Queries and Incremental Voronoi Diagrams
The Complexity of Diffuse Reflections in a Simple Polygon
Counting Proportions of Sets: Expressive Power with Almost Order
Efficient Approximate Dictionary Look-Up for Long Words over Small Alphabets
Relations Among Notions of Security for Identity Based Encryption Schemes
Optimally Adaptive Integration of Univariate Lipschitz Functions
Classical Computability and Fuzzy Turing Machines
An Optimal Algorithm for the Continuous/Discrete Weighted 2-Center Problem in Trees
An Algorithm for a Generalized Maximum Subsequence Problem
Random Bichromatic Matchings
Eliminating Cycles in the Discrete Torus
On Behalf of the Seller and Society: Bicriteria Mechanisms for Unit-Demand Auctions
Pattern Matching Statistics on Correlated Sources
Robust Model-Checking of Linear-Time Properties in Timed Automata
The Computational Complexity of the Parallel Knock-Out Problem
Reconfigurations in Graphs and Grids
${\mathcal C}$-Varieties, Actions and Wreath Product
Local Construction of Planar Spanners in Unit Disk Graphs with Irregular Transmission Ranges
An Efficient Approximation Algorithm for Point Pattern Matching Under Noise
Oblivious Medians Via Online Bidding
Efficient Computation of the Relative Entropy of Probabilistic Automata
A Parallel Algorithm for Finding All Successive Minimal Maximum Subsequences
De Dictionariis Dynamicis Pauco Spatio Utentibus
Customized Newspaper Broadcast: Data Broadcast with Dependencies
On Minimum {\itshape k}-Modal Partitions of Permutations
Two Birds with One Stone: The Best of Branchwidth and Treewidth with One Algorithm
Maximizing Throughput in Queueing Networks with Limited Flexibility
Network Flow Spanners
Finding All Minimal Infrequent Multi-dimensional Intervals
Cut Problems in Graphs with a Budget Constraint
Lower Bounds for Clear Transmissions in Radio Networks
Asynchronous Behavior of Double-Quiescent Elementary Cellular Automata
Lower Bounds for Geometric Diameter Problems
Connected Treewidth and Connected Graph Searching
A Faster Algorithm for Finding Maximum Independent Sets in Sparse Graphs
The Committee Decision Problem
Common Deadline Lazy Bureaucrat Scheduling Revisited
Approximate Sorting
Stochastic Covering and Adaptivity
Algorithms for Modular Counting of Roots of Multivariate Polynomials
Hardness Amplification Via Space-Efficient Direct Products
The Online Freeze-Tag Problem
I/O-Efficient Algorithms on Near-Planar Graphs
Minimal Split Completions of Graphs
Design and Analysis of Online Batching Systems
Competitive Analysis of Scheduling Algorithms for Aggregated Links
A 4-Approximation Algorithm for Guarding 1.5-Dimensional Terrains
On Sampling in Higher-Dimensional Peer-to-Peer Systems
Mobile Agent Rendezvous in a Synchronous Torus
Randomly Colouring Graphs with Girth Five and Large Maximum Degree
Packing Dicycle Covers in Planar Graphs with No {\itshape K}5--{\itshape e} Minor
Sharp Estimates for the Main Parameters of the Euclid Algorithm
Position-Restricted Substring Searching
Rectilinear Approximation of a Set of Points in the Plane
The Branch-Width of Circular-Arc Graphs
Minimal Eulerian Circuit in a Labeled Digraph
Speeding up Approximation Algorithms for NP-Hard Spanning Forest Problems by Multi-objective Optimization
{\sc RISOTTO}: Fast Extraction of Motifs with Mismatches
Minimum Cost Source Location Problems with Flow Requirements
Exponential Lower Bounds on the Space Complexity of OBDD-Based Graph Algorithms
Constructions of Approximately Mutually Unbiased Bases
Improved Exponential-Time Algorithms for Treewidth and Minimum Fill-In
Backmatter
Recommend Papers

LATIN 2006: Theoretical Informatics: 7th Latin American Symposium, Valdivia, Chile, March 20-24, 2006, Proceedings (Lecture Notes in Computer Science, 3887)
 9783540327554, 354032755X

  • 0 0 0
  • Like this paper and download? You can publish your own PDF file online for free in a few minutes! Sign Up
File loading please wait...
Citation preview

Lecture Notes in Computer Science Commenced Publication in 1973 Founding and Former Series Editors: Gerhard Goos, Juris Hartmanis, and Jan van Leeuwen

Editorial Board David Hutchison Lancaster University, UK Takeo Kanade Carnegie Mellon University, Pittsburgh, PA, USA Josef Kittler University of Surrey, Guildford, UK Jon M. Kleinberg Cornell University, Ithaca, NY, USA Friedemann Mattern ETH Zurich, Switzerland John C. Mitchell Stanford University, CA, USA Moni Naor Weizmann Institute of Science, Rehovot, Israel Oscar Nierstrasz University of Bern, Switzerland C. Pandu Rangan Indian Institute of Technology, Madras, India Bernhard Steffen University of Dortmund, Germany Madhu Sudan Massachusetts Institute of Technology, MA, USA Demetri Terzopoulos New York University, NY, USA Doug Tygar University of California, Berkeley, CA, USA Moshe Y. Vardi Rice University, Houston, TX, USA Gerhard Weikum Max-Planck Institute of Computer Science, Saarbruecken, Germany

3887

José R. Correa Alejandro Hevia Marcos Kiwi (Eds.)

LATIN 2006: Theoretical Informatics 7th Latin American Symposium Valdivia, Chile, March 20-24, 2006 Proceedings

13

Volume Editors José R. Correa Universidad Adolfo Ibáñez, School of Business Av. Presidente Errázuiz 3485, Las Condes, Santiago, Chile url: http://www.jcorea.uai.cl Alejandro Hevia Universidad de Chile, Department of Computer Science Blanco Encalada 2120, Santiago Centro, Santiago, Chile url: http://www.dcc.uchile.cl/∼ahevia Marcos Kiwi Universidad de Chile Center for Mathematical Modelling and Department of Mathematical Engineering Blanco Encalada 2120, Santiago Centro, Santiago, Chile url: http://www.dim.uchile.cl/∼mkiwi

Library of Congress Control Number: Applied for CR Subject Classification (1998): F.2, F.1, E.1, E.3, G.2, G.1, I.3.5, F.3, F.4 LNCS Sublibrary: SL 1 – Theoretical Computer Science and General Issues ISSN ISBN-10 ISBN-13

0302-9743 3-540-32755-X Springer Berlin Heidelberg New York 978-3-540-32755-4 Springer Berlin Heidelberg New York

This work is subject to copyright. All rights are reserved, whether the whole or part of the material is concerned, specifically the rights of translation, reprinting, re-use of illustrations, recitation, broadcasting, reproduction on microfilms or in any other way, and storage in data banks. Duplication of this publication or parts thereof is permitted only under the provisions of the German Copyright Law of September 9, 1965, in its current version, and permission for use must always be obtained from Springer. Violations are liable to prosecution under the German Copyright Law. Springer is a part of Springer Science+Business Media springer.com © Springer-Verlag Berlin Heidelberg 2006 Printed in Germany Typesetting: Camera-ready by author, data conversion by Scientific Publishing Services, Chennai, India Printed on acid-free paper SPIN: 11682462 06/3142 543210

Preface

This volume contains the papers accepted for publication at LATIN 2006, the 7th Latin American Theoretical Informatics Symposium held in Valdivia, Chile, March 20-24, 2006. The LATIN series of conferences presents recent results in theoretical computer science. It was launched in 1992 to foster the interaction between the Latin American community and computer scientists around the world. LATIN 2006 was the seventh of a series, after S˜ ao Paulo, Brazil (1992); Valparaiso, Chile (1995); Campinas, Brazil (1998), Punta del Este, Uruguay (2000), Canc´ un, Mexico (2002), and Buenos Aires, Argentina (2004). In response to the call for papers, a record number of 224 submissions were received. The Program Committee accepted 66 papers in order to meet the goal of having five days of talks with no parallel sessions. Therefore, many good papers could not be accepted. The Program Committee met electronically from October 25 to November 10, 2005. The selection of papers was based on originality, quality, and relevance to theoretical computer science. It is expected that most of these papers will appear in a more complete and polished form in scientific journals in the future. In addition to the contributed papers, this volume contains the abstracts of seven invited plenary talks given at the conference by Ricardo Baeza-Yates, Anne Condon, Ferran Hurtado, R. Ravi, Madhu Sudan, Sergio Verd´ u and Avi Wigderson. The Program Committee thanks all authors of submitted manuscripts for their support of LATIN, and the many colleagues listed in pages VIII-X, who helped reviewing the submissions. The LATIN proceedings have been published by Springer since the first edition. We are grateful to Springer for their continuous support.

January 2006

Jos´e R. Correa (Co-organizer) Alejandro Hevia (Co-organizer) Marcos Kiwi (Program Chair) LATIN 2006

The Conference

Program Committee Argimiro Arratia Quesada Marcelo Arenas Jose Correa Luc Devroye Guillermo Dur´ an Antonio Fern´ andez David Fern´ andez-Baca Afonso Ferreira Fedor Fomin Joachim von zur Gathen Cristina Gomes Fernandes Mordecai Golin Alejandro Hevia Marcos Kiwi (Chair) Hugo Krawczyk Eduardo Laber Martin Matamala Jiˇr´ı Matouˇsek Gonzalo Navarro Daniel Panario Andreas Schulz Jacques Sakarovitch Amin Shokrollahi Mona Singh Howard Straubing Shang-Hua Teng Luca Trevisan Jorge Urrutia Umesh Vazirani Santosh Vempala Alfredo Viola Yoshiko Wakabayashi

U. Valladolid, Spain U. Toronto, Canada U. Adolfo Ib´ an ˜ ez, Chile McGill U., Canada U. Chile, Chile / U. Bs. Aires, Argentina U. Rey Juan Carlos, Spain Iowa State U., USA CNRS & COST Office, France U. Bergen, Norway Bonn U., Germany U. de S˜ao Paulo, Brazil Hong Kong UST, Hong Kong, China U. Chile, Chile U. Chile, Chile IBM Research, USA PUC-Rio, Brazil U. Chile, Chile Charles U., Czech Republic U. Chile, Chile Carleton U., Canada MIT, USA CNRS/ENST, France EPFL, Switzerland Princeton U., USA Boston College, USA Boston U., USA UC Berkeley, USA UNAM, Mexico UC Berkeley, USA MIT, USA U. de la Rep´ ublica, Uruguay U. de S˜ ao Paulo, Brazil

Organizing Committee Jos´e Rafael Correa (Co-organizer) Alejandro Hevia (Co-organizer) Michael Vielhaber (Local liaison)

U. Adolfo Ib´ an ˜ ez, Chile U. Chile, Chile U. Austral, Chile

VIII

Organization

Sterring Committee Ricardo Baeza-Yates (Chair) Martin Farach-Colton Gaston Gonnet Yoshiharu Kohayakawa Sergio Rajsbaum Gadiel Seroussi

U. Chile, Chile Rutgers U., USA ETH Z¨ urich, Switzerland U. de S˜ ao Paulo, Brazil U. Nacional Aut´onoma M´exico, Mexico MSRI, USA

de

Plenary Speakers Ricardo Baeza-Yates Anne Condon Ferran Hurtado R. Ravi Madhu Sudan Sergio Verd´ u Avi Wigderson

U. Chile, Chile U. British Columbia, Canada U. Polit`ecnica de Catalunya, Spain Carnegie Mellon U., USA MIT, USA Princeton U., USA Institute for Advanced Study, USA

Sponsors Centro Latinoamericano de Estudios en Inform´ atica (CLEI) Centro de Modelamiento Matem´ atico (CMM), U. Chile CONICYT via grant Anillo en Redes International Federation for Information Processing (IFIP)

Referees Scott Aaronson Laila El Aimani Tatsuya Akutsu Liliana Alcon Sergio Alvarez Andris Ambainis Arvind Arasu Sergio Ar´evalo Ofer Arieli Esther Arkin Sunil Arya Albert Atserias Jean-Michel Autebert

Nikhil Bansal Denilson Barbosa Pablo Barcelo Amotz Barnoy David Mix Barrington Michael Bender Leopoldo Bertossi Gustavo Betarte Silvia Bianchi Ernesto G. Birgin Hans L. Bodlaender Alexandra Boldyreva Maria Luisa Bonet

Flavia Bonomo Claudson Bornstein Endre Boros Jeremie Bourdon Xavier Boyen Richard Brent Hajo Broersma Andrei Bulatov Samuel Buss Benjamin Bustos Leizhen Cai Alberto Caprara Rafael Carrasco

Organization

Carlos Castillo Marcia Cerioli Melissa Chase Fr´ed´eric Chataigner Kamalika Chaudhuri Edgar Ch´ avez Bernard Chazelle Siu-Wing Cheng Joseph Cheriyan Francis Y.L. Chin Vicent Cholvi Rezaul Alam Chowdhury Ferdinando Cicalese Julien Clement Peter Clote Anne Condon William Cook Graham Cormode Derek Corneil Bruno Courcelle Ricardo Dahab Josep Diaz Volker Diekert Ajit Diwan Yevgeniy Dodis Gerhard Dorfer Feodor Dragan Michael Drmota Georges Dupret Mariana Escalante Daniel Espinoza Luerbio Faria Uriel Feige Jose Fernandez Esteban Feuerstein Jiri Fiala Celina Figueiredo Samuel Fiorini Marc Fischlin David Flores Pierre Fraigniaud Matthew Franklin Michael J. Freedman Armin F¨ ugenschuh Henryk Fuks

Martin F¨ urer Peter Gacs David Gamarnik Theo Garefalakis Paul Gastin Mark Giesbrecht Michael Goldwasser Jovan Golic Martin Golumbic Marcos Goycoolea Eduardo Grampin Lov K. Grover Irene Guessarian Gregory Gutin Andr´ as Gy´arf´ as Michel Habib Yosub Han Pierre Hansen Meng He Pinar Heggernes Maurice Herlihy Miki Hermann Celina Herrera Shlomo Hershkop Susan Hohenberger Stefan Hougardy Ferran Hurtado Hsien-Kuei Hwang John Iacono Costas Iliopoulos Neil Immerman Nicole Immorlica Daniel Alejandro Jaume Emmanuel Jeandel Shaoquan Jiang Charanjit Jutla Bala Kalyanasundaram Sampath Kannan Juhani Karhumaki Marek Karpinski Eike Kiltz Kwangjo Kim Ekkehard Koehler Yoshiharu Kohayakawa Michael K¨ohler

Jochen K¨ onemann Daniel Kral Dieter Kratsch Manfred Kufleitner Peter Kugel Alberto Lafuente Alair Pereira Do Lago Tanja Lange John Langford Federico Lecumberry Orlando Lee Xiang-Yang Li Leonid Libkin Min Chih Lin Claudia Linhares Andrea Lodi Sylvain Lombardy Alex Lopez-Ortiz Florent Madelaine Fr´ed´eric Magniez Mohammad Mahdian Veli Makinen Arnaldo Mandel Yishay Mansour Javier Marenco Maurice Margenstern Alvaro Mart´ın Conrado Mart´ınez F´ abio Viduani Martinez Guillermo Matera Jacques Mazoyer Ross McConnell Luis A. A. Meira Donatella Merlini Ramgopal Mettu Gatis Midrijanis Peter Bro Miltersen Lorenz Minder Joseph Mitchell F. K. Miyazawa Manal Mohamed Cristopher Moore Ruggero Morselli J. Ian Munro

IX

X

Organization

Ashwin Nayak Sergio Nesmachnow Gregory Neven Alantha Newman Dejan Nickovic Harald Niederreiter Johannes Nowak Juha Nurmonen Michael N¨ usken Jos´e Oncina-Carratal´ a Alan Oppenheim Gianpaolo Oriolo Luis Ortiz Rafail Ostrovsky Alberto Pardo Alvaro Pardo Rodrigo Paredes Ojas Parekh David Parkes Rene Peralta Gonzalo Perera Dominique Perrin Artur Pessoa Benny Pinkas Valentin Polishchuk Helmut Prodinger F´ abio Protti Mihai Prunescu Artem Pyatkin Charles Rackoff Rajmohan Rajaraman

Rajeev Raman Michael Rao Ivan Rapaport Andr´e Raspaud Bruce Reed Omer Reingold Pablo Rey Franco Robledo Virg´ınia M. Rodrigues Andrea Rodriguez Frank Ruskey Alexander Russell Michel Saks Juan Jose Salazar Matias Salibian-Barrera Nicolas Schabanel Jay Sethuraman Ronen Shaltiel David Shmoys Jamshid Shokrollahi Igor Shparlinksi Martin Skutella Adam Smith Siang W. Song Gregory Sorkin Jeremy Spinrad Rob van Stee Lorna Stewart Nicolas E. Stier-Moses Paul K. Stockmeyer Maxim Sviridenko

Mario Szegedy Jayme L. Szwarcfiter Arie Tamir Dimitrios Thilikos Christopher Thraves Ioan Todinca Vilmar Trevisan Akio Tsuneda Mario Valencia-Pabon Gabriel Valiente Brigitte Vallee Eric Vigoda Jes´ us Villadangos Yngve Villanger Jan Vondr´ ak Gustavo Vulcano Michael Wagner John Watrous Marcelo Weinberger Renato Werneck Jiri Wiedermann Thomas Wilke Ryan Williams Erik Winfree Prudence Wong Derik Wood Moti Yung Paula Zabala Fangguo Zhang Yong Zhang

Table of Contents

Keynotes Algorithmic Challenges in Web Search Engines Ricardo Baeza-Yates . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1

RNA Molecules: Glimpses Through an Algorithmic Lens Anne Condon . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

8

Squares Ferran Hurtado . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

11

Matching Based Augmentations for Approximating Connectivity Problems R. Ravi . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

13

Modelling Errors and Recovery for Communication Madhu Sudan . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

25

Lossless Data Compression Via Error Correction Sergio Verd´ u ..................................................

26

The Power and Weakness of Randomness in Computation Avi Wigderson . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

28

Regular Contributions A New GCD Algorithm for Quadratic Number Rings with Unique Factorization Saurabh Agarwal, Gudmund Skovbjerg Frandsen . . . . . . . . . . . . . . . . . . .

30

On Clusters in Markov Chains Nir Ailon, Steve Chien, Cynthia Dwork . . . . . . . . . . . . . . . . . . . . . . . . . . .

43

An Architecture for Provably Secure Computation Mikl´ os Ajtai, Cynthia Dwork, Larry Stockmeyer . . . . . . . . . . . . . . . . . . .

56

Scoring Matrices That Induce Metrics on Sequences El´ oi Ara´ ujo, Jos´e Soares . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

68

XII

Table of Contents

Data Structures for Halfplane Proximity Queries and Incremental Voronoi Diagrams Boris Aronov, Prosenjit Bose, Erik D. Demaine, Joachim Gudmundsson, John Iacono, Stefan Langerman, Michiel Smid . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

80

The Complexity of Diffuse Reflections in a Simple Polygon Boris Aronov, Alan R. Davis, John Iacono, Albert Siu Cheong Yu . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

93

Counting Proportions of Sets: Expressive Power with Almost Order Argimiro Arratia, Carlos E. Ortiz . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 Efficient Approximate Dictionary Look-Up for Long Words over Small Alphabets Abdullah N. Arslan . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 Relations Among Notions of Security for Identity Based Encryption Schemes Nuttapong Attrapadung, Yang Cui, David Galindo, Goichiro Hanaoka, Ichiro Hasuo, Hideki Imai, Kanta Matsuura, Peng Yang, Rui Zhang . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 Optimally Adaptive Integration of Univariate Lipschitz Functions Ilya Baran, Erik D. Demaine, Dmitriy A. Katz . . . . . . . . . . . . . . . . . . . . 142 Classical Computability and Fuzzy Turing Machines Benjam´ın Ren´e Callejas Bedregal, Santiago Figueira . . . . . . . . . . . . . . . 154 An Optimal Algorithm for the Continuous/Discrete Weighted 2-Center Problem in Trees Boaz Ben-Moshe, Binay Bhattacharya, Qiaosheng Shi . . . . . . . . . . . . . . 166 An Algorithm for a Generalized Maximum Subsequence Problem Thorsten Bernholt, Thomas Hofmeister . . . . . . . . . . . . . . . . . . . . . . . . . . . 178 Random Bichromatic Matchings Nayantara Bhatnagar, Dana Randall, Vijay V. Vazirani, Eric Vigoda . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 190 Eliminating Cycles in the Discrete Torus B´ela Bollob´ as, Guy Kindler, Imre Leader, Ryan O’Donnell . . . . . . . . . . 202 On Behalf of the Seller and Society: Bicriteria Mechanisms for Unit-Demand Auctions Claudson Bornstein, Eduardo Sany Laber, Marcelo Mas . . . . . . . . . . . . . 211

Table of Contents

XIII

Pattern Matching Statistics on Correlated Sources J´er´emie Bourdon, Brigitte Vall´ee . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224 Robust Model-Checking of Linear-Time Properties in Timed Automata Patricia Bouyer, Nicolas Markey, Pierre-Alain Reynier . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 238 The Computational Complexity of the Parallel Knock-Out Problem Hajo Broersma, Matthew Johnson, Dani¨el Paulusma, Iain A. Stewart . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250 Reconfigurations in Graphs and Grids Gruia Calinescu, Adrian Dumitrescu, J´ anos Pach . . . . . . . . . . . . . . . . . . 262 C-Varieties, Actions and Wreath Product Laura Chaubard . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 274 Local Construction of Planar Spanners in Unit Disk Graphs with Irregular Transmission Ranges Edgar Ch´ avez, Stefan Dobrev, Evangelos Kranakis, Jaroslav Opatrny, Ladislav Stacho, Jorge Urrutia . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 286 An Efficient Approximation Algorithm for Point Pattern Matching Under Noise Vicky Choi, Navin Goyal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 298 Oblivious Medians Via Online Bidding Marek Chrobak, Claire Kenyon, John Noga, Neal E. Young . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 311 Efficient Computation of the Relative Entropy of Probabilistic Automata Corinna Cortes, Mehryar Mohri, Ashish Rastogi, Michael D. Riley . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323 A Parallel Algorithm for Finding All Successive Minimal Maximum Subsequences Ho-Kwok Dai, Hung-Chi Su . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 337 De Dictionariis Dynamicis Pauco Spatio Utentibus (lat. On Dynamic Dictionaries Using Little Space) Erik D. Demaine, Friedhelm Meyer auf der Heide, Rasmus Pagh, Mihai Pˇ atra¸scu . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 349 Customized Newspaper Broadcast: Data Broadcast with Dependencies Sandeep Dey, Nicolas Schabanel . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 362

XIV

Table of Contents

On Minimum k-Modal Partitions of Permutations Gabriele Di Stefano, Stefan Krause, Marco E. L¨ ubbecke, Uwe T. Zimmermann . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374 Two Birds with One Stone: The Best of Branchwidth and Treewidth with One Algorithm Frederic Dorn, Jan Arne Telle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386 Maximizing Throughput in Queueing Networks with Limited Flexibility Douglas G. Down, George Karakostas . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398 Network Flow Spanners Feodor F. Dragan, Chenyu Yan . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 410 Finding All Minimal Infrequent Multi-dimensional Intervals Khaled M. Elbassioni . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 423 Cut Problems in Graphs with a Budget Constraint Roee Engelberg, Jochen K¨ onemann, Stefano Leonardi, Joseph (Seffi) Naor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 435 Lower Bounds for Clear Transmissions in Radio Networks Mart´ın Farach-Colton, Rohan J. Fernandes, Miguel A. Mosteiro . . . . . 447 Asynchronous Behavior of Double-Quiescent Elementary Cellular Automata ´ Nazim Fat`es, Damien Regnault, Nicolas Schabanel, Eric Thierry . . . . . 455 Lower Bounds for Geometric Diameter Problems Herv´e Fournier, Antoine Vigneron . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 467 Connected Treewidth and Connected Graph Searching Pierre Fraigniaud, Nicolas Nisse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 479 A Faster Algorithm for Finding Maximum Independent Sets in Sparse Graphs Martin F¨ urer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491 The Committee Decision Problem Eli Gafni, Sergio Rajsbaum, Michel Raynal, Corentin Travers . . . . . . . 502 Common Deadline Lazy Bureaucrat Scheduling Revisited Ling Gai, Guochuan Zhang . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 515 Approximate Sorting Joachim Giesen, Eva Schuberth, Miloˇs Stojakovi´c . . . . . . . . . . . . . . . . . . 524

Table of Contents

XV

Stochastic Covering and Adaptivity Michel Goemans, Jan Vondr´ ak . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 532 Algorithms for Modular Counting of Roots of Multivariate Polynomials Parikshit Gopalan, Venkatesan Guruswami, Richard J. Lipton . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 544 Hardness Amplification Via Space-Efficient Direct Products Venkatesan Guruswami, Valentine Kabanets . . . . . . . . . . . . . . . . . . . . . . . 556 The Online Freeze-Tag Problem Mikael Hammar, Bengt J. Nilsson, Mia Persson . . . . . . . . . . . . . . . . . . . 569 I/O-Efficient Algorithms on Near-Planar Graphs Herman Haverkort, Laura Toma . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 580 Minimal Split Completions of Graphs Pinar Heggernes, Federico Mancini . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 592 Design and Analysis of Online Batching Systems Regant Y.S. Hung, Hing-Fung Ting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 605 Competitive Analysis of Scheduling Algorithms for Aggregated Links Wojciech Jawor, Marek Chrobak, Christoph D¨ urr . . . . . . . . . . . . . . . . . . 617 A 4-Approximation Algorithm for Guarding 1.5-Dimensional Terrains James King . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629 On Sampling in Higher-Dimensional Peer-to-Peer Systems Goran Konjevod, Andr´ea W. Richa, Donglin Xia . . . . . . . . . . . . . . . . . . 641 Mobile Agent Rendezvous in a Synchronous Torus Evangelos Kranakis, Danny Krizanc, Euripides Markou . . . . . . . . . . . . . 653 Randomly Colouring Graphs with Girth Five and Large Maximum Degree Lap Chi Lau, Michael Molloy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665 Packing Dicycle Covers in Planar Graphs with No K5 − e Minor Orlando Lee, Aaron Williams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 677 Sharp Estimates for the Main Parameters of the Euclid Algorithm Lo¨ıck Lhote, Brigitte Vall´ee . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 689 Position-Restricted Substring Searching Veli M¨ akinen, Gonzalo Navarro . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703

XVI

Table of Contents

Rectilinear Approximation of a Set of Points in the Plane Yan Mayster, Mario A. Lopez . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 715 The Branch-Width of Circular-Arc Graphs Fr´ed´eric Mazoit . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 727 Minimal Eulerian Circuit in a Labeled Digraph Eduardo Moreno, Mart´ın Matamala . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 737 Speeding Up Approximation Algorithms for NP-Hard Spanning Forest Problems by Multi-objective Optimization Frank Neumann, Marco Laumanns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745 RISOTTO: Fast Extraction of Motifs with Mismatches Nadia Pisanti, Alexandra M. Carvalho, Laurent Marsan, Marie-France Sagot . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 757 Minimum Cost Source Location Problems with Flow Requirements Mariko Sakashita, Kazuhisa Makino, Satoru Fujishige . . . . . . . . . . . . . . 769 Exponential Lower Bounds on the Space Complexity of OBDD-Based Graph Algorithms Daniel Sawitzki . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 781 Constructions of Approximately Mutually Unbiased Bases Igor E. Shparlinski, Arne Winterhof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 793 Improved Exponential-Time Algorithms for Treewidth and Minimum Fill-In Yngve Villanger . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 800 Author Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 813

Algorithmic Challenges in Web Search Engines Ricardo Baeza-Yates1,2 1

2

Center for Web Research, Dept. of Computer Science, Universidad de Chile, Santiago, Chile ICREA Professor, Dept. of Technology, Universitat Pompeu Fabra, Barcelona, Catalunya, Spain [email protected]

Abstract. In this paper we present the main algorithmic challenges that large Web search engines face today. These challenges are present in all the modules of a Web retrieval system, ranging from the gathering of the data to be indexed (crawling) to the selection and ordering of the answers to a query (searching and ranking). Most of the challenges are ultimately related to the quality of the answer or the efficiency in obtaining it, although some are relevant even to the existence of search engines: context based advertising.

1

Introduction

The Web is the largest public collection of data, and therefore, Web search has become one of the main challenges in the field of information retrieval. The complexity is not only due to its volume, but also because of its dynamics and heterogeneity. In addition, as search engines are (still) free as a consequence of Web advertising, the choice and placement of advertisements in the answer page is crucial to their survival. For all these reasons, we believe that Web retrieval is one of the main sources for interesting and challenging applied algorithmic problems. A Web search engine has basically four software modules around an index, as shown in Figure 1. We know detail this simplified software architecture. The Crawling module brings new or updated pages to the Indexing module. The Indexing module creates a compact searchable index with key preprocessed information for page ranking. The Searching module, using the index, finds a ranked answer to a stream of queries from remote users. Finally, the Answering module creates the answer page and places the right advertisement related to a query. Each of these modules presents a different set of challenges, which motivate this paper. The next sections summarize the main algorithmic challenges of the software modules mentioned before, using simplified versions of each problem so that we can (more or less) formalize them. We include recent results, although the bibliography is by no means exhaustive. The final section mentions additional problems related to the Web1 . 1

Disclaimer: The choice of problems is biased to the preferences of the author who is now at Yahoo! Research.

J.R. Correa, A. Hevia, and M. Kiwi (Eds.): LATIN 2006, LNCS 3887, pp. 1–7, 2006. c Springer-Verlag Berlin Heidelberg 2006 

2

R. Baeza-Yates

Fig. 1. Simplified software architecture of a Web search engine

2

Crawling

A crawler is a software module that gathers Web pages to create an index, typically using a parallel and distributed architecture. In practice a crawler never stops as the Web keeps growing, but we simplify the problem by limiting the crawling time and scope. Given a set of Web sites with their bandwidth W , a period of time T , a set of politeness rules P [14], a set of resources R (computers, crawler bandwidth, etc.), and three functions V , Q and F (over a set of pages), a crawler has to bring a collection of pages C ∈ W , achieving the following four main goals: – – – –

maximize the volume V (C) of the pages, maximize the quality Q(C) of the pages, minimize the freshness F (C) of the pages, and maximize the use of R while satisfying the politeness rules P .

Part of the problem is how to define the three optimization functions and how to combine them to obtain the best possible collection C that should imply the best possible index. Each possible choice brings new problems. It is out of the scope of this paper to present an exhaustive list of possibilities, but here are a few examples: – V could be the number of different pages or the total number of bytes of text. The former brings another problem, detection of duplicated pages, while the latter raises the question of how many bytes (or percentage) are needed to have a good textual description of a page; – Q could be based on the distribution of words or the link structure of the Web (e.g. Pagerank [17]); and – F could be the absolute time difference between the time when the page was crawled and the last modification time. However this raises the problem of measuring such function, as we cannot know its value until the end. Hence, usually F is an estimation of the freshness.

Algorithmic Challenges in Web Search Engines

3

Considering the above functions as given, we could define an overall function to optimize such as f (C) = αV (C) + βQ(C) − γF (C), for some α, β, γ > 0. Regarding R and P , we can simplify R as the maximum bandwidth available, B, and P as a minimum number of seconds, s, the wait time between any two page requests to the same site. Notice that being polite to Web sites opposes the goal of using all the bandwidth available at the search engine side. Then we have a formal scheduling problem: find a sequence of requests for complete pages2 at given times (p1 , t1 ), · · · , (pn , tn ), n > 0, such that we maximize f (C) (C = ∪ni=1 pi ), satisfying the following: – Crawling period: tn − t1 ≤ T , – Politeness: |ti − tj | ≥ s for any pair of pages pi , pj that belong to the same Web site, and – Overall bandwidth: for any time τ (t1 ≤ τ ≤ tn ), b(Wτ ) ≤ B; where b(Wτ ) is the bandwidth of the set of active requests at a given time τ . A request (pi , ti ) is active if pi belongs to site wi ∈ Wτ , and if ti ≤ τ ≤ ti + size(pi )/b(wi ) (b(wi ) is the bandwidth of the site wi 3 ). In the open scope case, finding new pages implies to find modified pages having new links, and that implies wasting time revisiting known pages. Hence, freshness opposes volume, given that wasted time cannot be used to crawl new pages. But, paradoxically, the number of new pages only increase by wasting that time. Another practical issue is dynamic pages, which can be unbounded. Several heuristics have been used, from breadth-first to ordering pages based on quality. Recently, a strategy based on Web site sizes (largest sites first, LSF) has shown to be competitive with strategies that use more information [3]. One problem is that there is no standard benchmark to compare crawling strategies, given that we would need the same Internet location for all the experiments.

3

Indexing and Searching

These two modules are interdependent, as the search time and the quality of the ranking will depend on the information stored in the index. Hence we present them as one integrated problem. The best index for searching words up to now is an inverted index, one of the oldest data structures [1]. An inverted index basically consists of a set of unique words (vocabulary) and a set of corresponding lists of pages where each word occurs. However, better solutions may exist, in particular given the new conditions imposed by the Web: smaller but distributed indexes, frequent updates, and fast answer time, to name a few. The index has to contain pre-calculated information that will be useful when ranking the answer. This information depends on the document similarity model 2

3

In practice could be partial pages, and in that case we also have to add to the request how many bytes to bring after we know the total size of the page. We assume that the bandwidth for each site is constant, which in practice is not true.

4

R. Baeza-Yates

used by the search engine and the query language used (they are not independent). Examples would be the vector model [1] that uses information on term distribution or a link based model such as Pagerank [17] or global Authorities [12]. Nowadays, search engines use many sources of information for ranking: text content, link structure, search engine usage, etc. One of the main problems is the evaluation of the quality of the ranking, as we do not know which are the best answers. Current evaluation techniques are based on click-through analysis (that is, how people click on the answer pages). Another restriction is the current Web volume, which implies that the index must be distributed across many machines. It also implies a partial evaluation of the query to give a fast response (in addition, as people on average looks at two answer pages, does not make sense to do additional work). If we add to this the current query volume, we need a parallel processing of the queries to increase the concurrency level of the overall retrieval system. Hence, we have a variant of the dictionary problem: design a dynamic data structure (the index) that, given a maximum space available M , achieves at least a throughput T of queries per second, satisfying the following requirements: – any query must be solved using only the index (accessing secondary memory or a remote page is too slow); – the index should be easy to distribute across many processors without making the search more difficult; – the index should be easy to update frequently, maintaining its quality and speed; and – the query can be partially evaluated to find the best B ranked answers. Almost all the requirements just mentioned imply some amount of extra information to be stored in the index. Hence, compression is an important element of current solutions. This can be approached from two directions: design a compressed searchable index, or, design a compression technique that allows fast searching [15]. In both cases, being able to search without the need of decompressing the index, improves the search time. In fact, these is one of the few cases where we can do faster search by using less space. Fast querying implies other interesting subproblems, such as fast computation of set operations (e.g. see [5]). Another of the most interesting subproblems are distributed indexes. For inverted indexes, there are two ways to distribute the index: – document partitioning splits the document collection in pieces and builds one partial index for each piece. Searching is achieved by merging partial answers from the processors that store the partial indexes. – term partitioning splits the vocabulary in pieces, and each processor holds one subset of the index, and hence, one part of a global inverted index. Searching is achieved by merging the answers for every word in the query. Current search engines use the former, as the main problem of the latter is that of building and maintaining a global index. However, new ideas may change in the future this choice, as term partitioning allows higher concurrency.

Algorithmic Challenges in Web Search Engines

4

5

Advertising

Deciding the right advertising is one of the main tasks of the Answering module. Context based advertising comes in two flavors: advertising shown after a query in the answer page of a search engine (e.g. Google’s AdWords) or the advertising shown in a syndicated page (e.g. Yahoo’s ContentMatch). The differences are the data available for matching the advertising: in one case a query and all its attributes, and in the second case the content of a page and the referrer information of the visitor to that page (which in some cases could include a query); and the number of places available (around ten in the former, two or three in the latter). In Figure 2 we present an example for the first case. Advertisers pay for each visit to a given site (this is called pay per click, PPC), so the search engine is interested in maximizing the future income. Clicks depend on the position of the ad, so the advertisers should be ranked. However, the choice of advertisers it is not as simple as choosing the ones that match and that would pay more for a click, as some advertisers are more clicked than others. Also, we cannot use click frequency as a definite rule, because advertisers that have had more exposition time, would have more clicks and new advertisers would then never appear, without having the chance to become popular and profitable. Then we have an on-line problem: given a set of evidences E regarding a page p, and a database of advertisers A, find a ranked subset a ∈ A that maximizes the expected income from clicks in p, such that |a| = n, where n is the number of places available in p.

Fig. 2. Example of keyword-based advertisement for the query “hotel valdivia”

6

R. Baeza-Yates

Notice that we do not separate the matching part (find advertisers that satisfy the evidence) from the ranking part (find advertisers that will maximize the future income), so E includes the data available for the match, past history, etc. There are few results for this problem. For example [18] explores the matching part, while [9, 16] the placement part. Another problem in the Answering module is the fast generation of the text summary (snippets) for each result in the answer page.

5

Concluding Remarks

We have briefly surveyed the algorithmic challenges of Web search. Solutions to these theoretical problems can help to find solutions to real ones, as in practice the problems are much more complex. For example, currently the Web has many forms of spam, including content, links and usage. So, finding the bad guys (e.g. pages that have misleading content, links that are used just to improve the ranking of the linked page, or clicks that come from malicious software agents) is an interesting dynamic problem. A possible easier solution could be to help the good guys, that is, the Web sites that do have good content and good links. Still, we have the problem of recognizing the good from the bad, which is related to the problem of information trust in the Semantic Web. This area is now called adversarial information retrieval(e.g. [11]). When we search we know what we are looking for. However, in the Web there could be interesting answers waiting to be discovered or interesting usage patterns that could help to improve a search engine. These are two examples of Web mining [10], a field that is still in its infancy. One interesting case is queries and the user actions after a query, or query mining [6]. When people formulate queries and click on answers, they are giving away “semantic information” for free, and with the current volume of queries per day in a search engine, the potential of this data is still unknown. For example, it could be used for better index design, better ranking, query optimization, query recommendation [4], generation of pseudo-semantic resources, Web site design [7], to name a few. We are currently building a platform to formalize Web mining tasks [8]. Finally, advertising is related to two newer fields: social networks [19] and Internet economy [20]. The intersection of these two fields will sparkle many interesting problems such as query incentive networks [13] or auction pricing.

Acknowledgments We thank the helpful comments of David Nettleton.

References 1. R. Baeza-Yates and B. Ribeiro-Neto, Modern Information Retrieval, AddisonWesley, England, 513 pages, 1999. 2. R. Baeza-Yates. Information Retrieval in the Web: beyond current search engines, Int. Journal of Approximate Reasoning 34 (2-3), 97-104, 2003.

Algorithmic Challenges in Web Search Engines

7

3. R. Baeza-Yates, C. Castillo, M. Marin, A. Rodriguez. Crawling a Country: Better Strategies than Breadth-First for Page Ordering. In WWW 2005, Industrial Track, ACM Press, Chiba, Japan, May 2005. 4. Ricardo A. Baeza-Yates, Carlos A. Hurtado, Marcelo Mendoza. Query Recommendation Using Query Logs in Search Engines, Current Trends in Database Technology - EDBT 2004 Workshops, Workshop on Clustering Information over the Web, Heraklion, Crete, Greece, March 14-18, 2004, Revised Selected Papers. LNCS 3268, Springer, p. 588-596. 5. R. Baeza-Yates, A Fast Set Intersection Algorithm for Sorted Sequences, In 15th Combinatorial Pattern Matching 2004, LNCS, Springer, Istanbul, Turkey, July 2004. 6. Ricardo Baeza-Yates. Applications of Web Query Mining. In European Conference on Information Retrieval (ECIR’05), D. Losada, J. Fern´ andez-Luna (editors), Springer LNCS 3408, Santiago de Compostela, Spain, March 2005, 7–22. 7. Ricardo Baeza-Yates, Barbara Poblete, A Website Mining Model Centered on User Queries, European Web Mining Forum, B. Berendt et al, editors. October 2005, Oporto, Portugal, p. 3-15. 8. R. Baeza-Yates, A. Pereira and N. Ziviani, WIM: A Web Information Mining Model for the Web, In LA-WEB 2005, IEEE CS Press, Oct 2005, 233-241. 9. H. K. Bhargava and J. Feng. Paid placement strategies for internet search engines. In Proceedings of the eleventh international conference on World Wide Web, pages 117-123. ACM Press, 2002. 10. S. Chakrabarti. Mining the Web: Discovering knowledge from hypertext data, Morgan Kaufmann, 2003. 11. B. Davison, editor. Workshop on Adversarial Information Retrieval on the Web. URL: http://airweb.cse.lehigh.edu/2005/. Chiba, Japan, May 2005. 12. J. Kleinberg. Authoritative sources in a hyperlinked environment. Journal of the ACM, 46(5):604–632, 1999. Preliminary version presented at SODA 1998. 13. J. Kleinberg and P. Raghavan. Query Incentive Networks. Proc. 46th IEEE Symposium on Foundations of Computer Science, 2005. 14. M. Koster. A standard for robot exclusion. http://www.robotstxt.org/wc/ exclusion.html, 1996. 15. V. Makinen and G. Navarro. Compressed Full Text Indexes. Technical Report TR/DCC-2005-7, Dept. of Computer Science, University of Chile, June 2005. Available at http://pizzachili.dcc.uchile.cl/biblio.html 16. S. Nicholson, T. Sierra, U.Y. Eseryel, J.H. Park, P. Barkow, E.J. Pozo, J. Ward. How Much of It is Real? Analysis of Paid Placement in Web Search Engine Results, JASIST, 2005. 17. L. Page, S. Brin, R. Motwani, and T. Winograd. The Pagerank citation algorithm: bringing order to the web. Technical report, Stanford Digital Library Technologies Project, 1998. 18. B. Ribeiro-Neto, M. Cristo, P. Golgher, and E. Silva de Moura. Impedance coupling in content-targeted advertising. In Proceedings of the 28th Annual international ACM SIGIR Conference on Research and Development in information Retrieval (Salvador, Brazil, August 15 - 19, 2005). SIGIR ’05. ACM Press, New York, NY, 496-503, 2005. 19. Barry Wellman. Computer Networks As Social Networks, Science 293 (5537), p. 2031-2034, 2001. 20. A.C-C. Yao, editor. First Workshop on Internet and Networks Economy, WINE 2005. URL: http://www.cs.cityu.edu.hk/ wine2005/, Hong-Kong, 2005.

RNA Molecules: Glimpses Through an Algorithmic Lens Anne Condon The Department of Computer Science, U. British Columbia, Canada [email protected]

Dubbed the “architects of eukaryotic complexity” [8], RNA molecules are increasingly in the spotlight, in recognition of the important catalytic and regulatory roles they play in our cells and their promise in therapeutics. Our goal is to describe the ways in which algorithms can help shape our understanding of RNA structure and function. Computational means for prediction of the structure of RNA and DNA molecules – collectively known as nucleic acids – are invaluable in determining the functions of molecules in the cell. Structure prediction problems, as well as the inverse problem of designing nucleic acids with specific structural properties, also arise in biological research aimed at creating new catalysts and biosensors, in nanotechnology, and in efforts to recreate an RNA world that may have preceded modern life [11]. Put simply, a DNA or RNA molecule is a sequence of units, called bases, over a four-letter alphabet. Prediction of nucleic acid structure is easier than prediction of protein structure, because the primary forces that determine nucleic acid structure are pairings (bonds) between individual bases of the molecule, with each base in at most one pair. This set of base pairs is called the secondary structure of the molecule. One premise is that, of the exponentially many possibilities, an RNA molecule folds into that secondary structure which has minimum free energy (MFE) [7]. Finding the MFE structure for a given RNA molecule is NP-hard [6]. However, the range of structures that arise in nature is relatively limited, making it feasible to find MFE predictions of almost all naturally-occurring secondary structures in polynomial time [10]. There is still much potential to advance the state of the art in MFE secondary structure prediction of nucleic acids. – The algorithm of Rivas and Eddy [10] is very general in terms of types of structures that it can predict [3], but with Θ(n6 ) running time is limited to relatively short inputs. One challenge is to find the sweet spot in the tradeoff between algorithmic generality and efficiency. A concrete goal is to find MFE “kissing hairpin” structures in less than O(n6 ) time. – MFE prediction algorithms can only be as good as their underlying energy models. While experimental wet-lab work has provided hundreds of highquality parameters for use in RNA structure prediction [7], some model features, including multi-loops and pseudoknots, have been parameterized in J.R. Correa, A. Hevia, and M. Kiwi (Eds.): LATIN 2006, LNCS 3887, pp. 8–10, 2006. c Springer-Verlag Berlin Heidelberg 2006 

RNA Molecules: Glimpses Through an Algorithmic Lens

9

a somewhat ad-hoc way, and indeed the model features themselves may have been chosen in part for algorithmic efficiency, rather than accuracy. Thus, we believe that a reassessment of the energy model, informed by physical principles, known nucleic acid structures, and algorithmic complexity, should be fruitful in improving the quality of prediction algorithms. Stepping back from our focus on MFE prediction from a single sequence, we note that there are many other interesting computational problems relating to nucleic acid structure prediction. Minimum free energy prediction of complexes of two or more molecules is an important goal that arises, for example, in determining the targets of anti-sense RNA’s [1]. Moreover, partition function prediction provides additional useful information, including base pairing probabilities [5]. In a different direction, the premise that molecules fold into their MFE structures may be false for some structures, when “kinetic traps” - low energy structures with no low-energy paths to a MFE structure - exist, or when folding occurs co-transcriptionally [9]. Thus, also important are alternative structure prediction approaches [9] and efficient simulation of folding kinetics [12, 13]. The problem of predicting three-dimensional structure is largely unsolved. Finally, although good heuristics have been developed for the design problem, which is to determine a sequence that folds into a given input secondary structure [2, 4], its computational complexity is still open.

References 1. C. Alkan, E. Karako¸c, J. Nadeau, S.C. Sahinalp, and K. Zhang. “RNA-RNA interaction prediction and antisense RNA target search”, Proc. Ninth Annual Intl. Conf. on Research in Computational Molecular Biology (RECOMB), 152–171, 2005. 2. M. Andronescu, A.P. Fejes, F. Hutter, H.H. Hoos, and A. Condon. “A new algorithm for RNA secondary structure design”, J. Mol. Biol., 336(3):607–624, 2004. 3. A. Condon, B. Davy, B. Rastegari, F. Tarrant, and S. Zhao. “Classifying RNA pseudoknotted structures”, Theoretical Computer Science, 320(1):35–50, 2004. 4. R.M. Dirks, M. Lin, E. Winfree and N.A. Pierce. “Paradigms for computational nucleic acid design”, Nucleic Acids Res., 32(4):1392–1403, 2004. 5. R.M. Dirks and N.A. Pierce. “A partition function algorithm for nucleic acid secondary structure including pseudoknots”, J. Comput. Chem., 24(13):1664–1677, 2004. 6. R.B. Lyngsø and C.N.S. Pedersen. “Pseudoknot prediction in energy based models”, J. Comput. Biol., 7:3409–427, 2000. 7. D.H. Mathews, J. Sabina, M. Zuker and D.H. Turner. “Expanded sequence dependence of thermodynamic parameters improves prediction of RNA secondary structure”, J. Mol. Biol., 288:911–940, 1999. 8. J.S. Mattick. “Non-coding RNAs: the architects of eukaryotic complexity”, EMBO (European Molecular Biology Organization) Reports, 2(11):986–991, 2001. 9. I.M. Meyer and I. Miklos. “Co-transcriptional folding is encoded within RNA genes”, BMC Molecular Biology, 5:10, 2004. 10. E. Rivas and S. R. Eddys. “A dynamic programming algorithm for RNA structure prediction including pseudoknots”, J. Mol. Biol., 225:2053–2068, 1999.

10

A. Condon

11. J.W. Szostak, D.P. Bartel, and L. Luisi. “Synthesizing life”, Nature, 409:387–389, 2001. 12. X. Tang, B. Kirkpatrick, S. Thomas, G. Song, and N.M. Amato. “Using Motion Planning to Study RNA Folding Kinetics”, J. Comput. Biol., 12(6):862–881, 2005. 13. A. Xayaphoummine, T. Bucher, and H. Isambert. “Kinefold web server for RNA/DNA folding path and structure prediction including pseudoknots and knots”, Nucleic Acid Res., 33:605–610, 2005.

Squares Ferran Hurtado Departament de Matem` atica Aplicada II, Universitat Polit`ecnica de Catalunya

In this talk we present several results and open problems having squares, the basic geometric entity, as a common thread. These results have been gathered from various papers; coauthors and precise references are given in the descriptions that follow. Disassembling Sets of Colored Squares Given a set of disjoint convex objects in the plane, it is well known that they can be moved to infinity without collision, one at a time, using only translations in the direction given by any vector v. If we have convex objects that have two different colors it is always possible to obtain two vectors v 1 and v 2 , forming possibly an infinitesimally small angle, such that each direction is used for one of the colors and the objects in their final far away positions are well separated, say by a line. An infinitesimally small angle is not quite satisfactory and it is natural to wonder whether a larger separating angle independent of n may always be obtained. This is precisely the problem we have considered in [5]: we study in which cases the separating angle can be guaranteed to be bounded by below, and consider also the similar problem for c colors. Somehow surprisingly, the shape of the objects happens to be a crucial issue; for example, any c-colored set of isothetic squares can be separated using separating vectors such that the angle between any two of them is at least π/(2c− 2), while for disks the situations is quite different, as the angle between separating vectors may be required to be arbitrarily small. We will discuss as well some algorithmic issues on these problems. Matching with Squares Let C be a class of geometric objects and let P be a point set with n elements p1 , . . . , pn in general position, n even. A C-matching of P is a set M = {C1 , . . . , Ck } of elements of C, such that every Ci contains exactly two elements of P . If all the elements of P belong to some Ci , M is called a perfect matching. If in addition all the elements of M are pairwise disjoint we say that the matching M is strong. If we define a graph GC (P ) in which the vertices are the elements of P , two of which are adjacent if there is an element of C containing them and no other element from P , a perfect matching in GC (P ) in the graph theory sense corresponds naturally with our definition of GC (P )-matchings. 

Research partially supported by Projects MCYT-FEDER BFM20033-0368 and Gen. Cat 2005SGR00692.

J.R. Correa, A. Hevia, and M. Kiwi (Eds.): LATIN 2006, LNCS 3887, pp. 11–12, 2006. c Springer-Verlag Berlin Heidelberg 2006 

12

F. Hurtado

If C is the set of all isothetic squares, M will be called a square-matching and the graph GC (P ) is the Delaunay graph for the L1 (or the L∞ ) metric. We prove that there is always a square-matching, which translates to these graphs to contain a perfect matching. In fact we have obtained a stronger result, namely that these graphs contain a Hamiltonian path, a question that remained unsolved since it was posed by Dillencourt [4], who proved the existence of perfect matchings in Delaunay triangulations for the Euclidean metric, i.e., that point sets always admit circle-matchings. On the other hand, the problem of deciding whether a point set admits a strong square-matching has recently been proved to be NP-hard [2]. The results we present were developed in [1], where this class of problems was introduced and studied on the light of geometric matchings. Tiny Squares We call a function I : Z2 → {0, 1} a binary image. We call the elements of Z2 pixels and we say that a pixel p is black (respectively, white) if I(p) = 1 (respectively, I(p) = 0). Taking the pixels as vertex set and the natural definitions of 4-neighbourhood and 8-neighbourhood the graphs G4 and G8 are obtained. For a given image I, its black and white pixels induce subgraphs that we denote B4 (I) and W4 (I), and B8 (I) and W8 (I) respectively. For a, b ∈ {4, 8} we say that an image I is Ba , Wb -connected if the graphs Ba (I) and Wb (I) are each connected, that is, each has a single connected component. In the paper [3] we consider for both graphs a local modification operation on binary images in which a black pixel p and a white pixel q can interchange their colours when they are neighbours, and we prove that, for any (a, b) ∈ {(4, 4), (4, 8), (8, 4), (8, 8)}, any two Ba , Wb -connected images I and J each with n black pixels can be converted into the other with a sequence of O(n2 ) 8-local interchanges if (a, b) ∈ {(4, 8), (8, 4), (8, 8)} and O(n4 ) 8-local interchanges if (a, b) = (4, 4).

References ´ 1. Bernardo Abrego, Esther Arkin, Silvia Fern´ andez, Ferran Hurtado, Mikio Kano, Joseph Mitchell and Jorge Urrutia Matching points with geometric objects, Proc. Japan Conf. on Discrete and Computational Geometry 2004, LNCS Springer-Verlag 3742, pp.1-15, 2005. 2. Sergey Bereg, Nikolaus Mutsanas and Alexander Wolff, Matching points with rectangles and squares, Proc. 32nd Int. Conf. on Current Trends in Theory and Practice of Computer Science, SOFSEM’06, January 2006, Merin, Czech Republic. 3. Prosenjit Bose, Vida Dujmovic, Ferran Hurtado and Pat Morin, ConnectivityPreserving Transformations of Binary Images, submitted for publication. 4. Michael Dillencourt, Toughness and Delaunay Triangulations, Discrete and Computational Geometry, 5(6), 1990, 575-601. Preliminary version in Proc. of the 3rd Ann. Symposium on Computational Geometry, Waterloo, 1987, 186-194. 5. Ferran Hurtado and Mikio Kano, On the translational separation of colored convex objects, Proc. XI Encuentros de Geometr´ıa Computacional, Santander, 2005, pp. 225-230.

Matching Based Augmentations for Approximating Connectivity Problems R. Ravi Tepper School of Business, Carnegie Mellon University, Pittsburgh, PA 15213 [email protected] Abstract. We describe a very simple idea for designing approximation algorithms for connectivity problems: For a spanning tree problem, the idea is to start with the empty set of edges, and add matching paths between pairs of components in the current graph that have desirable properties in terms of the objective function of the spanning tree problem being solved. Such matching augment the solution by reducing the number of connected components to roughly half their original number, resulting in a logarithmic number of such matching iterations. A logarithmic performance ratio results for the problem by appropriately bounding the contribution of each matching to the objective function by that of an optimal solution. In this survey, we trace the initial application of these ideas to traveling salesperson problems through a simple tree pairing observation down to more sophisticated applications for buy-at-bulk type network design problems.

1

Introduction

Approximation algorithms have been traditionally designed and taught on a problem-by-problem basis; Surveys (e.g., [6]) and recent courses and books (e.g., [13]) have approached the area in this way by mainly classifying key results based on a problem-specific basis. As the field matures to provide a rich variety of results, commonalities can be identified to highlight key techniques that become repeatedly useful. In this survey, we point to one such extremely simple technique that we term MBA, an acronym for Matching Based Augmentation. The two salient features that determine the applicability of the method are that the problem at hand must be a connectivity problem where one tries to connect up various demands (either among themselves or to a common root) in a network, and that the optimal solution can be used to identify an appropriate polynomial-time solvable augmenting subproblem that is a variant of matching. Since the method proceeds by finding such matching iteratively and adding them to augment the solution, the approximation ratio is typically bounded by the number of iterations of the process; Furthermore, since the cost paid by the augmentation 

Supported in part by NSF grant CCF-0430751 and ITR grant CCR-0122581 (The ALADDIN project).

J.R. Correa, A. Hevia, and M. Kiwi (Eds.): LATIN 2006, LNCS 3887, pp. 13–24, 2006. c Springer-Verlag Berlin Heidelberg 2006 

14

R. Ravi

in each iteration is bounded with respect to the optimal, the method of proof of the performance ratio is primal-based, without relying on any new (lower) bounds on the (minimization) problem to argue the guarantee. Finally, since each augmentation proceeds by matching up components of the current solution, the number of iterations before there is one single component and hence a feasible solution, is logarithmic in the number of demand points that need to be connected. This explains why most approximation algorithms based on this method have logarithmic performance ratio. We have structured this survey chronologically by describing the applications of the method in the order of their first (typically conference or technical report) publication. In this order, the classic paper of Christofides [2] is the first paper of the sequence to contain most features of the MBA idea: the missing idea is the iterative augmentation. The ATSP approximation of Frieze et al. [3] uses the MBA idea in it complete form to obtain a logarithmic approximation for metric ATSPs. We review these two results in the next section. In the following section, we trace our own work in a series of papers [7, 10, 12, 8, 11] that use this idea in various contexts for NP-hard undirected spanning tree problems. In the next section, we review some more sophisticated uses of the method to solve generalizations of basic connectivity problems so as to route flow under concave cost functions [9, 1, 5]. We